# Observable for research - (advanced) statistics

**URL:** <https://talk.observablehq.com/t/observable-for-research-advanced-statistics/4538>\
**Category:** Uncategorized\
**Created:** [January 27, 2021, 5:18pm UTC](https://talk.observablehq.com/t/observable-for-research-advanced-statistics/4538 "2021-01-27T17:18:46Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![chrispahm](https://yyz2.discourse-cdn.com/flex030/user_avatar/talk.observablehq.com/chrispahm/32/4203_2.png) [@chrispahm](https://talk.observablehq.com/u/chrispahm)\
**Post date:** [January 27, 2021, 5:18pm UTC](https://talk.observablehq.com/t/observable-for-research-advanced-statistics/4538/1 "2021-01-27T17:18:46Z")

</div>

Hey Observables folks 👋

Here’s something that I’d like to discuss with you:  
Are you planning on using Observable notebooks for research / teaching? Or are you already doing so? If not, what is keeping you from doing it?

Here’s my current stance on it:  
**Pros**  
One idea I absolutely love about Observables is that users do not need to  
install any software in order to edit and run a notebook.  
You can just send a link to supervisors, students, reviewers, fellow researchers,  
as well as your parents, and they can have a look at your notebook and play  
around with it.

Another great benefit is that due to the reactive approach, the execution order does not matter in an Observable notebook. Just last month the JetBrains Datalore Team examined 10.000.000 Jupyter notebooks on Github, and found that ~36% would not run right away due to out-of-order execution [source](https://blog.jetbrains.com/datalore/2020/12/17/we-downloaded-10-000-000-jupyter-notebooks-from-github-this-is-what-we-learned/).

I think the combination of the two points not only shows the potential of Observables to overcome the technical part of the [reproducibility crisis](https://en.wikipedia.org/wiki/Replication_crisis), but also highlights the lower entrance barrier to statistics / data science / programming in general.

**Cons**  
Unfortunately, there still seems to be a lack of a well maintained (advanced) statistics library in the JS ecosystem, especially when it comes to doing regression analysis.  
While I think [simple-statistics](https://simplestatistics.org),  
[jstats](http://jstat.github.io), and [ml.js](https://github.com/mljs/ml#regression))  
are great libraries, they are missing the functionality to perform (still relatively basic) regressions on (1) multiple-linear models (2) linear models including higher powers of one or more predictor variables (3) models with interaction terms (4) models including weights ([see here for a non-extensive list](http://mezeylab.cb.bscb.cornell.edu/labmembers/documents/supplement%205%20-%20multiple%20regression.pdf)).

The lack of these functions currently pushes me (and probably others) to using Python/R/Matlab for the actual data analysis, and uploading the results to Observable afterwards for visualization.  
Unfortunately, this way we’re losing the benefit of having the data, analysis, and visualization all in one place, where students / fellow researchers / reviewers etc. can easily play around with the data / method.

**Synopsis**  
While I am aware that it is possible to [execute R from an Observable notebook](https://observablehq.com/@bryangingechen/hello-r-on-opencpu), I still think that there is a need for an actual JS (or WebAssembly) implementation of the core functions (IMHO something similar to [`lm`](http://madrury.github.io/jekyll/update/statistics/2016/07/20/lm-in-R.html) would be needed the most).

Right now we are teaching students R, but I sometimes think teaching JS/Observable might get more people interested in the subject as it also allows them to produce Websites / d3 visualizations.

Really interested in your thoughts and ideas here!

---

<div class="post-metadata">

**Author:** ![tomlarkworthy](https://yyz2.discourse-cdn.com/flex030/user_avatar/talk.observablehq.com/tomlarkworthy/32/5940_2.png) [@tomlarkworthy](https://talk.observablehq.com/u/tomlarkworthy)\
**Post date:** [January 27, 2021, 6:12pm UTC](https://talk.observablehq.com/t/observable-for-research-advanced-statistics/4538/2 "2021-01-27T18:12:09Z")

</div>

Yes I agree. We had a kinda similar chat on the Zulip server about how numerical computing holds back JS. ([https://d3js.zulipchat.com/#narrow/stream/273726-observable/topic/numerical.20computing.20on.20the.20web](https://d3js.zulipchat.com/#narrow/stream/273726-observable/topic/numerical.20computing.20on.20the.20web))

R, Python and Matlab link to Fortan’s actively maintained BLAS, LAPACK. This is the missing foundation that stats is built on IMHO.

Python has finalizers which means it can make an ergonomic wrapper around the memory managed  
primitives (numpy). I think JS did not have that until recently with WeakRef and FinalizationRegistry. So basically it was impossible to ergonomically attach to the de factor linear libraries ([GitHub - likr/emlapack: BLAS / LAPACK for JavaScript](https://github.com/likr/emlapack)). You would have to call “free()” in JS land.

Anyway, we were literally just chatting about how it might actually be possible now (not in a cross browser way). But AFAIK no-one has combined FinalizationRegistry and WASM for LAPACK so it be a pretty researchy thing to do and probably a fail.

But yeah, definitely agree the inability to do linear programming, eigenvectors, lasso regression or a trillion other off-the-shelf algorithms almost mandates using something else or sticking to primitive analysis. Its very annoying.

---

<div class="post-metadata">

**Author:** ![mcmcclur](https://yyz2.discourse-cdn.com/flex030/user_avatar/talk.observablehq.com/mcmcclur/32/2364_2.png) [@mcmcclur](https://talk.observablehq.com/u/mcmcclur)\
**Post date:** [January 27, 2021, 8:10pm UTC](https://talk.observablehq.com/t/observable-for-research-advanced-statistics/4538/3 "2021-01-27T20:10:05Z")

</div>

I use Observable to teach _elementary_ statistics regularly but probably not in quite the way you are thinking. It sounds like you are thinking of using Javascript in an Observable notebook kinda like one might use R in RStudio or Python in Jupyter. That’s not what I do. Rather, I build interactive illustrations of the main ideas with Observable. If you’re familiar with R, it’s somewhat analogous to the way one might build an illustration with Shiny.

You can tak a look at my [class web page from last semester](https://marksmath.org/classes/Fall2020Stat185/) to see what I’m talking about. Illustrations built with Observable are really littered all over. Having said that, you’ll also notice that I still use Python within Google Colab as the computational environment.

On the other hand, I have built some elementary computational tools for intro stats, like [this T-Distribution calculator](https://marksmath.org/classes/Fall2020Stat185/t_calculators.html) which is built on [this Observable notebook](https://observablehq.com/@mcmcclur/the-t-distribution). I’ve thought of expanding those greatly so that students could easily import or upload CSV files and run elementary statistical tests. I don’t think it would be too hard to build most of what one needs in intro stats. Not sure if I’ll ever get to that or not.

---

<div class="post-metadata">

**Author:** ![chrispahm](https://yyz2.discourse-cdn.com/flex030/user_avatar/talk.observablehq.com/chrispahm/32/4203_2.png) [@chrispahm](https://talk.observablehq.com/u/chrispahm)\
**Post date:** [January 27, 2021, 10:10pm UTC](https://talk.observablehq.com/t/observable-for-research-advanced-statistics/4538/4 "2021-01-27T22:10:28Z")

</div>

Thanks a lot for you insights! I remember checking out the elmlapack repo back when I first read this blog post on how R actually fits a linear model ([A Deep Dive Into How R Fits a Linear Model](http://madrury.github.io/jekyll/update/statistics/2016/07/20/lm-in-R.html)), until then I didn’t know that it’s actually running Fortran in the background! Also, just stumbled upon this pure JS re-write: [GitHub - R-js/blasjs: Pure Javascript manually written implementation of BLAS, Many numerical software applications use BLAS computations, including Armadillo, LAPACK, LINPACK, GNU Octave, Mathematica, MATLAB, NumPy, R, and Julia.](https://github.com/R-js/blasjs)

Back then I was naïvely thinking I could re-write the underlying R-code (until the point where R-calls the C function `Cdqrls`, which then calls Fortran), and turn it into a native module for Node.js… However, I quickly realised that the task is beyond my capabilities.

However, what are your thoughts on expanding existing libraries like `simple-statistics` to a point where a simple ANOVA or the previously mentioned (non-basic) linear regressions are possible? The other day I started implementing a function for weighted linear regressions, but adding support for multiple predictors is on another level!

---

<div class="post-metadata">

**Author:** ![chrispahm](https://yyz2.discourse-cdn.com/flex030/user_avatar/talk.observablehq.com/chrispahm/32/4203_2.png) [@chrispahm](https://talk.observablehq.com/u/chrispahm)\
**Post date:** [January 27, 2021, 10:33pm UTC](https://talk.observablehq.com/t/observable-for-research-advanced-statistics/4538/5 "2021-01-27T22:33:00Z")

</div>

That’s great to hear that you use Observables in class already! With @tomlarkworthy’s response in mind, I was thinking that it might be an intermediate step to extent a library auch as `simple-statistics`, before someone starts writing a NumPy/SciPy equivalent in JS.

Just as you said, it would be really nice for students to be able to upload a CSV file, and then perform an ANOVA / (Multiple) Linear-Regression analysis right in the notebook. While they might not be the most performant alternatives, I guess there already are packages for most of the required functions for this. However, I couldn’t find something even remotely close to R’s own `lm` yet, especially when it comes to multiple linear regressions.

---

<div class="post-metadata">

**Author:** ![mcmcclur](https://yyz2.discourse-cdn.com/flex030/user_avatar/talk.observablehq.com/mcmcclur/32/2364_2.png) [@mcmcclur](https://talk.observablehq.com/u/mcmcclur)\
**Post date:** [January 27, 2021, 11:02pm UTC](https://talk.observablehq.com/t/observable-for-research-advanced-statistics/4538/6 "2021-01-27T23:02:05Z")

</div>

> [@chrispahm](#):
>
> However, what are your thoughts on expanding existing libraries like `simple-statistics`

I’m afraid I have to offer a word of caution on `simple-statistics`; I’m just a bit leery of the quality under the hood.

I wrote a simple calculator using `ss.inverseErrorFunction` for students to compute some percentile ranks and they just weren’t quite getting the right answers. I traced the problem down to the implementation which, as I recall used a something like an interpolation of some sparsely pre-computed values from a table. I ended up quickly writing an implementation that used Newton’s method to invert the normal PDF. I’m sure that’s not particularly efficient but it worked.

Since then, I’ve moved to [jstat](https://github.com/jstat/jstat) and have been quite happy with it.

I think this sheds light on the general problem. Building quality numerical tools is a long road. In this particular case, we have a truly outstanding developer but who’s not a statistician. It’s also not clear how much vetting the library has been through. You need a community of folks who are good programmers and good scientists.

* * *

The idea of implementing BLAS in Javascript is not fathomable to me. 😕

---

<div class="post-metadata">

**Author:** ![jrus](https://yyz2.discourse-cdn.com/flex030/user_avatar/talk.observablehq.com/jrus/32/586_2.png) [@jrus](https://talk.observablehq.com/u/jrus)\
**Post date:** [January 27, 2021, 11:19pm UTC](https://talk.observablehq.com/t/observable-for-research-advanced-statistics/4538/7 "2021-01-27T23:19:35Z")

</div>

I use Observable for what I would call “research” projects, but it doesn’t involve statistical analysis. (And most of the actual ‘research’ part is done with pen and paper. There is also sometimes some Photoshop, Python, and Matlab used as tools along the way.)

As for BLAS: Maybe someone can compile Fortran to wasm effectively. A web search turns up [GitHub - StarGate01/Full-Stack-Fortran: Fortran to WebAssembly](https://github.com/StarGate01/Full-Stack-Fortran)

---

<div class="post-metadata">

**Author:** ![jrus](https://yyz2.discourse-cdn.com/flex030/user_avatar/talk.observablehq.com/jrus/32/586_2.png) [@jrus](https://talk.observablehq.com/u/jrus)\
**Post date:** [January 27, 2021, 11:34pm UTC](https://talk.observablehq.com/t/observable-for-research-advanced-statistics/4538/8 "2021-01-27T23:34:13Z")

</div>

> [@mcmcclur](#):
>
> I wrote a simple calculator using `ss.inverseErrorFunction` for students to compute some percentile ranks and they just weren’t quite getting the right answers. I traced the problem down to the implementation which, as I recall used a something like an interpolation of some sparsely pre-computed values from a table.

The implementation hasn’t changed in 6 years, and this does not seem like an accurate description. The approximation used is Winitzki’s. The reference is this google hosted PDF file: [erf-approx.pdf - Google Drive](https://drive.google.com/file/d/0B2Mt7luZYBrwZlctV3A3eF82VGM/view)

> <https://github.com/simple-statistics/simple-statistics/blob/29e3f3b7644d1d99a63ff6cf093d20b2a91ef781/src/inverse_error_function.js#L7-L19>

There are a number of possible approximations to inverse erf, depending on your needs. It is plausible that this approximation is not accurate enough for your students needs though. @mbostock made a notebook describing this and a couple others: [Error Function / Mike Bostock | Observable](https://observablehq.com/@mbostock/error-function)

---

<div class="post-metadata">

**Author:** ![mcmcclur](https://yyz2.discourse-cdn.com/flex030/user_avatar/talk.observablehq.com/mcmcclur/32/2364_2.png) [@mcmcclur](https://talk.observablehq.com/u/mcmcclur)\
**Post date:** [January 28, 2021, 12:43am UTC](https://talk.observablehq.com/t/observable-for-research-advanced-statistics/4538/9 "2021-01-28T00:43:14Z")

</div>

> [@jrus](#):
>
> The approximation used is Winitzki’s.

Thanks for the clarification. Of course, there are often trade-offs to be made in numerical computation and my _guess_ is that the code in Simple Statistics favors speed over precision. Whether that’s right or wrong, my comparisons with Mathematica indicate that

```
ss.inverseErrorFunction(x)

```

is off in the fourth decimal place for all x with |x|\>0.68. That just wasn’t sufficient for my purposes. Hence, my word of caution. 😐

---

<div class="post-metadata">

**Author:** ![jrus](https://yyz2.discourse-cdn.com/flex030/user_avatar/talk.observablehq.com/jrus/32/586_2.png) [@jrus](https://talk.observablehq.com/u/jrus)\
**Post date:** [January 28, 2021, 12:58am UTC](https://talk.observablehq.com/t/observable-for-research-advanced-statistics/4538/10 "2021-01-28T00:58:01Z")

</div>

I agree it would probably be better to default to something relatively accurate to at least 7 or 8 digits, or even try to compute something to full accuracy with an approximation available as an explicit alternative. Or if not, to make the limitations crystal clear in the docs.

They would probably take a patch of a more accurate approximation. This one seems to have been added in 2015 by a user without too much other demand or discussion, and hasn’t been mentioned since in github issues, so probably isn’t seeing much practical use. [Add inverses for error\_function and cumulative\_std\_normal\_probability by ericfischer · Pull Request #85 · simple-statistics/simple-statistics · GitHub](https://github.com/simple-statistics/simple-statistics/pull/85)

---

<div class="post-metadata">

**Author:** ![mbostock](https://yyz2.discourse-cdn.com/flex030/user_avatar/talk.observablehq.com/mbostock/32/9_2.png) [@mbostock](https://talk.observablehq.com/u/mbostock)\
**Post date:** [January 28, 2021, 2:37am UTC](https://talk.observablehq.com/t/observable-for-research-advanced-statistics/4538/11 "2021-01-28T02:37:43Z")

</div>

I think [https://stdlib.io/](https://stdlib.io/) has a BLAS package, if that helps?

[https://stdlib.io/docs/api/v0.0.90/@stdlib/blas](https://stdlib.io/docs/api/v0.0.90/@stdlib/blas)

One challenge with stdlib is that it’s currently in a monolithic 223 MB package, which does not make it easily consumable in a web browser. I suspect that there’s a way to repackage it, though.

There’s an example of elmapack on Observable, too:

> **[Hello matrix-eig, hello emlapack](https://observablehq.com/@fil/hello-matrix-eig)**
>
> emlapack is a BLAS/LAPACK package compiled with emscripten. In other words, it’s a complicated piece of software that does (hopefully) quick & accurate linear algebra computations. matrix-eig is a wrapper around emlapack that facilitates the...

---

<div class="post-metadata">

**Author:** ![jrus](https://yyz2.discourse-cdn.com/flex030/user_avatar/talk.observablehq.com/jrus/32/586_2.png) [@jrus](https://talk.observablehq.com/u/jrus)\
**Post date:** [January 28, 2021, 3:31am UTC](https://talk.observablehq.com/t/observable-for-research-advanced-statistics/4538/12 "2021-01-28T03:31:01Z")

</div>

> [@mbostock](#):
>
> I suspect that there’s a way to repackage it, though.

@rreusser has been working on this

> **[GitHub - stdlib-js/esm: ES module distribution for stdlib, a standard library...](https://github.com/stdlib-js/esm)**
>
> ES module distribution for stdlib, a standard library for JavaScript and Node.js. - GitHub - stdlib-js/esm: ES module distribution for stdlib, a standard library for JavaScript and Node.js.

---

<div class="post-metadata">

**Author:** ![mbostock](https://yyz2.discourse-cdn.com/flex030/user_avatar/talk.observablehq.com/mbostock/32/9_2.png) [@mbostock](https://talk.observablehq.com/u/mbostock)\
**Post date:** [January 28, 2021, 4:27am UTC](https://talk.observablehq.com/t/observable-for-research-advanced-statistics/4538/13 "2021-01-28T04:27:34Z")

</div>

Thanks, that seems to work (though it might be more efficient bundled):

```auto
blas = import("https://cdn.jsdelivr.net/npm/@stdlib/esm@0.0.3/blas.js") 

```

---

<div class="post-metadata">

**Author:** ![mootari](https://yyz2.discourse-cdn.com/flex030/user_avatar/talk.observablehq.com/mootari/32/581_2.png) [@mootari](https://talk.observablehq.com/u/mootari)\
**Post date:** [January 28, 2021, 10:07am UTC](https://talk.observablehq.com/t/observable-for-research-advanced-statistics/4538/14 "2021-01-28T10:07:08Z")

</div>

For what it’s worth, I started collecting discussions around stdlib-js and similar libraries here:

> **[Stdlib-js Survey](https://observablehq.com/@mootari/stdlib-survey)**
>
> Variants: "stdlib", "stdlibjs", "stdlib.js", "stdlib-js", "stdlib.io" Unrelated projects with similar names: stdlib.com / Github Org "stdlib" (Autocode Standard Library, serverless) npm package "stdlib" (Node/Sails modules) Kotlin stdlib-js general...

---

<div class="post-metadata">

**Author:** ![aaronkyle](https://yyz2.discourse-cdn.com/flex030/user_avatar/talk.observablehq.com/aaronkyle/32/8106_2.png) [@aaronkyle](https://talk.observablehq.com/u/aaronkyle)\
**Post date:** [January 28, 2021, 12:48pm UTC](https://talk.observablehq.com/t/observable-for-research-advanced-statistics/4538/15 "2021-01-28T12:48:44Z")

</div>

H folks,

This is a bit distant from the discussion, but in line with the topic heading of this thread, I thought it might be worthwhile flagging one major challenge of using JavaScript (in general) for advanced statistics, namely the breakdown of computational capacity for very large numbers:

> **[Weird](https://observablehq.com/@bumbeishvili/interesting)**

---

<div class="post-metadata">

**Author:** ![ds604](https://yyz2.discourse-cdn.com/flex030/user_avatar/talk.observablehq.com/ds604/32/625_2.png) [@ds604](https://talk.observablehq.com/u/ds604)\
**Post date:** [January 28, 2021, 6:20pm UTC](https://talk.observablehq.com/t/observable-for-research-advanced-statistics/4538/16 "2021-01-28T18:20:17Z")

</div>

Thanks for a concise breakdown of the technical barriers. For people who don’t follow Javascript-land that closely, but are kind of waiting for it to get there, it’s hard to get a sense for what exactly is the holdup, or to name specific technical issues that are holding things back (that might give some sense of a timeline for when the issues might be addressed).

The other side of the issue (what would researchers need for this to be viable) I guess would be the front-end (slice syntax and operator overloading for a R/Numpy/Matlab vectorized look), which might be where Observable could play a role. I recall seeing [NumCalc](http://numcalc.com/) a few years ago, and being pretty interested to see that kind of language extension (which i guess is kind of an alternative to operator overloading for the whole Javascript, something like that). I was wondering if Observable might at some point be using the cell separation to allow some of that syntax experimentation (for all the scientist/domain people who might be pouring into Javascript-land at some point, and want their home notation to work with).

(I was following what [Iodide](https://alpha.iodide.io/) was doing a while ago wrt some of these steps, but I think they’ve kind of slowed down.)

---

<div class="post-metadata">

**Author:** ![jrus](https://yyz2.discourse-cdn.com/flex030/user_avatar/talk.observablehq.com/jrus/32/586_2.png) [@jrus](https://talk.observablehq.com/u/jrus)\
**Post date:** [January 28, 2021, 7:47pm UTC](https://talk.observablehq.com/t/observable-for-research-advanced-statistics/4538/17 "2021-01-28T19:47:04Z")

</div>

Operator overloading would be super helpful. There is a proposal for it: [https://github.com/tc39/proposal-operator-overloading](https://github.com/tc39/proposal-operator-overloading). But I don’t expect it to happen any time soon.

---

<div class="post-metadata">

**Author:** ![dkirkby](https://yyz2.discourse-cdn.com/flex030/user_avatar/talk.observablehq.com/dkirkby/32/3775_2.png) [@dkirkby](https://talk.observablehq.com/u/dkirkby)\
**Post date:** [January 29, 2021, 1:17am UTC](https://talk.observablehq.com/t/observable-for-research-advanced-statistics/4538/18 "2021-01-29T01:17:44Z")

</div>

Like @mcmcclur, I often use observable to prepare interactive teaching materials (for undergraduate cosmology) but I don’t expose the underlying notebooks to the students. I would guess this a more common approach when students are not expected to be javascript literate. My student-facing materials are [here](https://dkirkby.github.io/cosmo-demo/) and the corresponding notebooks are in [this collection](https://observablehq.com/collection/@dkirkby/cosmology).

My use of observable for teaching is definitely limited by the availability of high-quality numerical libraries, but I have been slowly implementing the pieces I need (e.g. this [adaptive integrator](https://observablehq.com/@dkirkby/integrate?collection=@dkirkby/utilities)).

---

<div class="post-metadata">

**Author:** ![jrus](https://yyz2.discourse-cdn.com/flex030/user_avatar/talk.observablehq.com/jrus/32/586_2.png) [@jrus](https://talk.observablehq.com/u/jrus)\
**Post date:** [January 29, 2021, 2:28am UTC](https://talk.observablehq.com/t/observable-for-research-advanced-statistics/4538/19 "2021-01-29T02:28:19Z")

</div>

Hi David,

> [@dkirkby](#):
>
> I have been slowly implementing the pieces I need (e.g. this [adaptive integrator](https://observablehq.com/@dkirkby/integrate)).

You might find useful the handful of things I put at [@jrus/cheb](https://observablehq.com/@jrus/cheb) about 2.5 years ago. There’s a ton more that could conceivably go in there. If you want to chat about this, [stop by here](https://d3js.zulipchat.com/#narrow/stream/273730-math).

---

<div class="post-metadata">

**Author:** ![mcmcclur](https://yyz2.discourse-cdn.com/flex030/user_avatar/talk.observablehq.com/mcmcclur/32/2364_2.png) [@mcmcclur](https://talk.observablehq.com/u/mcmcclur)\
**Post date:** [January 29, 2021, 11:20am UTC](https://talk.observablehq.com/t/observable-for-research-advanced-statistics/4538/20 "2021-01-29T11:20:53Z")

</div>

> [@dkirkby](#):
>
> I don’t expose the underlying notebooks to the students

I certainly don’t either! While I share the implementations on this forum, the demos are embedded into class webpages on my site.

[Next page](https://talk.observablehq.com/t/observable-for-research-advanced-statistics/4538.md?page=2)
