ResearchOS/Wiki

Data Hub

Data Hub is a free, open-source GraphPad Prism alternative built into ResearchOS. It runs your statistics and your publication figures entirely in the browser, so your data never leaves your folder.

Data Hub. The navigator rail holds your tables, analyses, and figures; the main panel shows the data grid, the results sheet, or the figure you are working on.
Watch the app demos

What Data Hub is and why it exists

Prism and its peers are graphical front ends over math that is already settled and already available as proven open-source code. The value they sell is the workflow, not the statistics. Data Hub reproduces that workflow on top of battle-tested open-source libraries, so you get the guided analysis experience without the per-seat license and without sending your unpublished data to anyone.

Everything runs on your own machine, in the browser. There is no server doing the computation, so your raw numbers stay in your data folder the same way your notes and sequences do. The reason this matters is simple. Your data is yours, the math is open, and you can read the exact code behind every number Data Hub reports.

Data Hub lives at /datahub and is organized like the rest of ResearchOS. Documents are scoped to projects (collections) with optional subfolders, and they stand on their own rather than being born attached to a single experiment, so one workbook can pull together data from wherever it came from.

The three-pane navigator

The left rail is the navigator. A collection selector at the top scopes the workbook list to one project or opens it to everything at once. Below it, the Data Tables section lists your tables in a foldered tree. Two more sections, Results and Graphs, hold the analyses and figures that belong to the table you have open. Selecting a table opens its data grid; selecting an analysis opens its results sheet; selecting a figure opens the graph editor.

Table types

Like Prism, Data Hub shapes the table to the analysis you intend, so the grid and the available tests match how your data is actually laid out. New table offers eight archetypes.

A Column table is the starting point. Each column is a treatment group and each row is a replicate. The footer shows the mean, SD, SEM, and n for every group, recomputed live as you type, with no separate summarize step.

An XY table pairs an X column with one or more Y columns, one observation per row. This is the shape for dose-response curves, time courses, and standard curves, and it unlocks correlation, linear regression, and fitted curves.

A Grouped table records two factors at once. Each row is a level of the row factor, each column group is a level of the column factor, and replicate subcolumns sit under each group. That replication is what lets a two-way ANOVA estimate the interaction between the two factors.

A Survival table records time-to-event data. Each row is a subject with a follow-up time, an event indicator (1 if the event happened, 0 if the subject was censored), and an optional group label so arms can be compared with a Kaplan-Meier curve and the log-rank test.

A Contingency table holds counts in a grid of two categorical factors. It runs the chi-square test of independence, and for a 2x2 table it adds Fisher's exact test with the relative risk and the odds ratio.

A Nested table records technical replicates nested inside biological replicates, for example cells within a mouse or mice within a treatment. It runs the nested t-test for two top-level groups and the nested one-way ANOVA for three or more, so technical replicates are not treated as independent.

A Parts of whole table describes the composition of a single whole, a category label and a value per slice. It draws pie, donut, and 100-percent stacked-bar figures, each slice as its percent of the total.

An Info sheet is a documentation table. It holds notes and constants that describe a dataset so the context travels with the data. It runs no analysis and draws no figure.

New table. Pick the archetype that matches your data; the grid and the available analyses follow from it.

The data grid and live summary

The grid is where you enter raw replicates. You only ever enter the numbers once. The summary statistics, every analysis, and every figure all read from the same cells, so an edit to a single replicate flows through to the footer, the results sheet, and any graph immediately. There is no recalculation step to forget.

Transforms and derived tables

Sometimes the raw numbers are not the numbers you analyze. Data Hub can produce a new table that is computed from an existing one, so the cleaning step is explicit and repeatable rather than a one-off edit you cannot retrace. Five transforms ship: a general transform that applies a function to every value, normalize, transpose (swap rows and columns), remove-baseline (subtract a chosen baseline), and fraction-of-total (each value as its share of its column or row).

A transform creates a derived table that keeps a live link to its source. When you open the derived table, it recomputes from the source's current content rather than from a saved snapshot, so an edit to the original flows through automatically and the derived table can never go stale. If the source is deleted, the derived table shows a clear empty state instead of silently keeping old numbers.

Running an analysis

New analysis offers only the tests that are valid for the open table, so you never pick a test the data cannot support. The set is broad and grows with the table type.

From a Column table: unpaired and paired t-tests, one-way ANOVA with Tukey comparisons, and their rank-based counterparts (Mann-Whitney U, Wilcoxon signed-rank, Kruskal-Wallis); repeated-measures ANOVA and its mixed-model cousin (a random-intercept linear mixed model) for within-subject designs; multiple linear regression; and the Grubbs outlier test.

From an XY table: Pearson and Spearman correlation, linear regression, simple logistic regression, the ROC curve with AUC, dose-response curve fitting (4PL and 5PL) with model comparison by AICc and an extra-sum-of-squares F test, and global (shared-parameter) fitting across several curves at once.

From the design-specific tables: two-way ANOVA from a Grouped table; Kaplan-Meier survival with the log-rank and Gehan-Breslow-Wilcoxon tests plus Cox proportional-hazards regression from a Survival table; the chi-square and Fisher exact tests from a Contingency table; and the nested t-test and nested one-way ANOVA from a Nested table.

Every result leads with a plain-language verdict, a sentence that states the practical takeaway before the numbers. Does it differ, which way, by how much. The full statistics table follows, laid out the way a methods section reads it.

The results sheet. The plain-language verdict comes first, then the full statistics table, then the pairwise comparisons.

A tour of the analyses

The same results-sheet pattern carries every analysis, from the everyday comparisons to the pharmacology fits and the survival models.

Dose-response. A 4PL or 5PL logistic fit reports the EC50 with an asymmetric confidence interval, the Hill slope, and the plateaus, with model comparison by AICc.
Survival. Kaplan-Meier curves per arm with the log-rank and Gehan-Breslow-Wilcoxon tests, plus Cox proportional-hazards regression.
ROC and AUC. The full curve, the area under it with a confidence interval, and the optimal cut point by Youden's J.
Linear regression. Slope and intercept with their standard errors and confidence intervals, R-squared, and the residual standard error.
Multiple regression. Each predictor with its coefficient, confidence interval, standardized beta, and VIF, plus the overall model fit.
Contingency. The chi-square test on the count matrix, and for a 2x2 table Fisher's exact test with relative risk and odds ratio.
Repeated measures. Within-subject ANOVA with the Greenhouse-Geisser and Huynh-Feldt sphericity corrections, alongside a random-intercept mixed model.
Outliers. The Grubbs test screens each column on its own, flagging values with their per-step G statistic against the critical value.

Show the code

Every analysis carries a Show the code toggle that reveals the exact open-source Python (scipy, statsmodels, lifelines) that reproduces it, with your real group names and values baked in. Paste it into a notebook and you get the same numbers. This is the answer to the question a closed tool cannot answer, which is where a given number actually came from.

The same proof extends to the picture. Every figure carries its own code export, the matplotlib that redraws it from the same group names and values the on-screen figure used, so the plot is as reproducible as the statistic. Both halves answer the same question for the number and for the figure.

Data Hub also drafts the paragraphs a paper needs. From a finished analysis it writes a Methods sentence that names the test and cites the canonical reference for it, a Results sentence that reports the finding with the inline statistics from the engine, and a formatted reference list that includes the open-source software the engine computed with. The numbers come only from the engine; the phrasing and the citations are the curation.

The guided analysis wizard

Most bench scientists are not statisticians, and the most common analysis mistake is running a t-test or an ANOVA on data that breaks the test's assumptions. The guided wizard exists to prevent that. It asks a few plain questions about what you are comparing, then it does not just name a test. It checks the assumptions and shows you what it found.

The assumption Report Card checks normality (Shapiro-Wilk, per group) and equal variance (Brown-Forsythe across groups), reports each as a plain-language pass or fail with a one-line reason, and when an assumption fails it falls back to the matching rank-based test that does not need that assumption, telling you why it switched. The recommendation you see is therefore already assumption-aware.

The guided wizard. It recommends a test and shows the assumption Report Card, switching to a rank-based test automatically when an assumption fails.

Publication-quality graphs

New graph generates a figure from the open table. The figure is real SVG, so it stays an infinitely scalable vector for a paper and also exports as a crisp hi-DPI PNG for a slide, or copies straight to the clipboard to paste into a document.

A Column table makes a column scatter (every replicate as a point over the group mean, the default, because individual points show the real spread a reviewer wants to see) or a bar with SD or SEM error bars. An XY table makes a scatter with a fitted curve laid over it (a least-squares line, a four-parameter dose-response logistic, Michaelis-Menten, exponential decay or association, and more). A Grouped table makes a grouped bar chart, one cluster per row level and one bar per group. A Survival table makes a Kaplan-Meier step curve, one line per arm.

For a Column table you can also draw an estimation plot, the modern effect-size figure that shows the raw data alongside the bootstrap sampling distribution of the mean difference and its confidence interval, in the Gardner-Altman style for two groups and the Cumming style for three or more sharing one control. It shows the size of the effect rather than only a yes or no significance star, and the interval comes straight from the same validated bootstrap the results sheet reports.

An estimation plot. It shows the raw data and the bootstrap distribution of the effect with its confidence interval, not just a significance star.

Error bars come straight from the raw replicates, the same numbers the grid footer shows, so a figure of a table is always consistent with that table. Significance brackets are pulled from a stored ANOVA, so the right stars drop onto the figure with one toggle rather than being drawn by hand.

Color is its own studio. The Palette Studio replaces a single color dropdown with a browseable palette library filtered by how many series the plot has, a custom per-series mode, a generate-and-lock workflow, import from a coolors.co URL, and your own saved palettes, with the live figure as the preview as you choose.

Cleaning a figure is part of the same loop. Right-click a data cell to exclude it: the value stays visible and editable in the grid, but every analysis and every plot treats it as absent, so it drops out of the mean, the error bars, the dots, and the stored test at once. Nothing is deleted and nothing is hidden.

Importing data

You do not have to retype data you already have. Data Hub imports CSV files, pasted Excel ranges (with a transpose option for the common wrong-orientation case), and binary .xlsx workbooks. It detects the header row and the per-column types and maps the data onto a table. The honest limit, accepted up front, is that no free browser library can round-trip a native embedded Excel chart, so Data Hub reads the data and the formulas and re-plots natively rather than trying to preserve the original chart object.

Import. Paste from Excel or pick a CSV or .xlsx file; Data Hub detects the structure and previews the table before creating it.

Referencing Data Hub in notes and results

A table, an analysis, or a figure can be referenced from a note or a result. The Copy reference button writes a link that renders as a live chip wherever you paste it, and clicking the chip opens the table in Data Hub. A figure can also be dropped into a note as an image with the figure's Copy button (which puts a PNG on the clipboard) or its Export PNG path. The chip keeps the live link to the analysis; the image captures the figure as it looked.

Connection to the rest of the app

Data Hub documents participate in the same project and folder structure as notes, experiments, methods, and sequences, and they go through the same Trash flow with the same recovery window when deleted. The point of building the analysis surface into ResearchOS rather than leaving it in a separate paid application is that your data, your statistics, and your figures all live in one place that you own.

A few threads tie it to the rest of the app. The power and sample-size planner answers the design question before any data is collected, how many subjects you need to detect an effect, how much power a planned n gives you, or the smallest effect that n can reliably detect, running against the same engine the analyses use. BeakerBot can create a Data Hub table for you from pasted data in one step, detecting the columns and writing the table. And a Data Hub table can drive a layer in the phylogenetics Tree Studio, so a metadata table can render as a tip-aligned plot beside a tree.