# Working with large file- Framework or other options?

**URL:** <https://talk.observablehq.com/t/working-with-large-file-framework-or-other-options/8834>\
**Category:** Help\
**Created:** [February 21, 2024, 9:53pm UTC](https://talk.observablehq.com/t/working-with-large-file-framework-or-other-options/8834 "2024-02-21T21:53:38Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![kickout](https://yyz2.discourse-cdn.com/flex030/user_avatar/talk.observablehq.com/kickout/32/2541_2.png) [@kickout](https://talk.observablehq.com/u/kickout)\
**Post date:** [February 21, 2024, 9:53pm UTC](https://talk.observablehq.com/t/working-with-large-file-framework-or-other-options/8834/1 "2024-02-21T21:53:38Z")

</div>

The timing of this framework release is eery. I am trying to solve a problem where I have a very large dataset (1.5G .csv or 300MB .zip). I need to dynamically filter this dataset based on 2 or 3 user inputs and then summarizes over US counties, and then plot on a choropleth. Not complicated.

Whenever I load this local file in the notebook, it basically crashes the notebook (which makes sense). Ideally I’d do the filtering outside the browser using R and then pass the summarized data back to Observable, but then I lose interactivity b/c it’s constantly taking inputs from JS/Observable, re-rerunning, etc.

The filtering and aggregation are simple. I’m just counting unique rows that meet the criteria. It’s weather data, so to filter in R and aggregate (using data.table syntax):  
`dat[wind_speed >10 & relative_humidity >75,][,length(date_iso),by=countyState]`

The JS equivalent isn’t complicated so I’ll omit it, but is Framework the only way to work with data this size in a semi-interactive manner? Can JS filter _as_ is reads as file?

---

<div class="post-metadata">

**Author:** ![Fil](https://yyz2.discourse-cdn.com/flex030/user_avatar/talk.observablehq.com/fil/32/207_2.png) [@Fil](https://talk.observablehq.com/u/Fil)\
**Post date:** [February 21, 2024, 10:15pm UTC](https://talk.observablehq.com/t/working-with-large-file-framework-or-other-options/8834/2 "2024-02-21T22:15:13Z")

</div>

It might pay off to consider using the parquet format (see [Sylvain Lesage: &quot;You should use Parquet for your data. Example #54…&quot; - Mastodon](https://mastodon.social/@severo/111957633001467414)). If your queries are “simple” (_a_ and _b_, like described above), you could also try a “data cube” strategy where the data is pre-aggregated at a certain (small) precision step; alternatively, splitting the dataset into chunks that you wouldn’t need to load all at the same time.
