ResearchOS/Wiki

Correlation and regression

When both of your variables are measured numbers, you usually want to know whether they move together and whether you can predict one from the other. Correlation measures how tightly two things track. Regression fits an actual line you can read a slope off of and use to predict. Multiple regression extends that to several predictors at once. This page covers all three and the trap that sits underneath them.

What correlation measures

Correlation measures how consistently two variables rise and fall together. The common measure is Pearson's r, which runs from -1 to +1. An r near +1 means that when one goes up the other reliably goes up; near -1 means one goes up as the other goes down; near 0 means no straight-line relationship. The closer to the ends, the tighter the cloud of points hugs a line.

The Data Hub reports r, its 95% confidence interval, a p-value for whether the correlation differs from zero, and often r squared, the fraction of the variation in one variable that tracks with the other. An r of 0.7 gives an r squared of about 0.49, so roughly half the variation is shared.

Spearman correlation

When your data are ranks, scores on an ordinal scale, or visibly non-normal, Spearman's rho is the nonparametric alternative. It works exactly like Pearson's r, but on the rank-transformed data rather than the raw values. The result is a coefficient that runs from -1 to +1 with the same meaning, and the Data Hub reports it with the same 95% confidence interval (Fisher z approximation on the ranks) and a p-value. Because Spearman captures any monotone relationship, not only straight lines, it is less sensitive to outliers and does not assume that the relationship between the two variables is linear.

Fitting a line with simple linear regression

Where correlation gives a single number for tightness, linear regression fits the actual line and hands you its equation. It is the right tool when one variable plausibly drives the other and you want to predict or quantify the relationship, the absorbance you expect at a given concentration, the signal per unit of input.

The result reports the slope with its confidence interval and p-value, the intercept, and r squared for how well the line fits. The slope is the headline, it is how much the outcome changes for each one-unit change in the predictor, in real units. A slope whose confidence interval excludes zero is a relationship you can stand behind.

The slope is the headline, the change in the outcome for each one-unit change in the predictor, in real units. It comes with its standard error and 95 percent confidence interval, and the r-squared says how much of the variation the line accounts for.

A worked example

You plot fluorescence against protein concentration and fit a line. The slope is 1,250 units per microgram (95% CI 1,180 to 1,320, p <0.0001), with r squared = 0.98. You would write "fluorescence rose 1,250 units per microgram of protein (95% CI 1,180 to 1,320, p <0.0001), and the linear fit explained 98% of the variance (r squared = 0.98)." The tight interval and high r squared together say this is a clean, usable standard curve.

Multiple regression: several predictors at once

Multiple regression predicts one outcome from two or more predictors together. Its real value is that each predictor's effect is estimated while holding the others constant. If yield depends on both temperature and pH, multiple regression tells you the effect of temperature at a fixed pH, separating two influences that a one-at-a-time analysis would tangle together.

The result reports, for each predictor, a coefficientwith its confidence interval and p-value, plus an overall r squared and a model-level p-value. Read each coefficient as "the change in the outcome per unit of this predictor, with the other predictors held fixed." A predictor that mattered on its own can fall to non-significance here, which usually means another predictor was carrying the signal all along.

Each predictor gets its own coefficient, read as the change in the outcome per unit while the other predictors are held fixed. The standardized beta puts the slopes on a common scale, and the VIF column flags predictors that move together, where a value above about 5 to 10 means a coefficient is hard to trust.

The VIF column and multicollinearity

The multiple-regression result includes a VIF (variance inflation factor) for each predictor. The VIF for a given predictor is computed by regressing that predictor on all the other predictors and taking 1 / (1 minus that r-squared). A VIF of 1 means the predictor is completely uncorrelated with the others, and its coefficient is as precisely estimated as the data allow. A VIF above roughly 5 to 10 is a flag that the predictor is redundant with something else in the model, and the coefficient's confidence interval is wider and less stable than it looks in isolation. When two predictors always move together in your setup, the model cannot separate their individual effects, and the coefficients for both get wide, shaky intervals. The fix is more varied data or dropping a redundant predictor, not trusting a narrow-looking coefficient.

ResearchOS validates Pearson and Spearman correlation, simple regression, and multiple regression (including VIF) against scipy and statsmodels on the transparency page.