Repair: Regression-Based
On this page
Summary
A multi-linear regression repairs a damaged curve by predicting it from good curves in the same well, fitted over undamaged samples near the gap. The result is a prediction, not a measurement: it carries the uncertainty of the fit and inherits any problem in its inputs. The calculator evaluates a fitted two-predictor model for density and its approximate band.
Inputs and outputs
| Item | Units | |
|---|---|---|
| Input | Gamma ray | gAPI |
| Input | Neutron porosity | v/v |
| Input | Regression intercept | g/cm³ |
| Input | Regression coefficient, gamma ray | g/cm³ per gAPI |
| Input | Regression coefficient, neutron | g/cm³ per v/v |
| Input | Regression residual standard deviation | g/cm³ |
| Output | Regression-predicted density | g/cm³ |
| Output | Regression band, low | g/cm³ |
| Output | Regression band, high | g/cm³ |
Equations
For \(p\) predictor curves \(x_1, \ldots, x_p\) and the damaged curve \(y\), the model is linear:
The coefficients are the least-squares solution over the \(n\) training samples, with design matrix \(X\) (a column of ones and the predictors):
The fit quality is the coefficient of determination and the residual standard deviation, with \(p+1\) fitted parameters:
For the two-predictor density model of the calculator, with the gamma ray \(\GR\) and the neutron porosity \(\phiN\):
An approximate 95 percent band is
The band uses only the scatter of the training residuals. It does not include the uncertainty of the coefficients, the error of the predictors, or any change in the relation between the training samples and the gap. It is a lower bound on the uncertainty.
| Symbol | Variable | Units | Typical range |
|---|---|---|---|
| \(\mathrm{GR}\) | Gamma ray | gAPI | 10 to 250 |
| \(\phi_N\) | Neutron porosity | v/v | -0.02 to 0.60 |
| \(\beta_0\) | Regression intercept | g/cm³ | |
| \(\beta_1\) | Regression coefficient, gamma ray | g/cm³ per gAPI | |
| \(\beta_2\) | Regression coefficient, neutron | g/cm³ per v/v | |
| \(\sigma_{\mathrm{fit}}\) | Regression residual standard deviation | g/cm³ | 0.01 to 0.05 |
| \(\hat{\rho}_b\) | Regression-predicted density | g/cm³ | |
| \(\hat{\rho}_{b,\mathrm{lo}}\) | Regression band, low | g/cm³ | |
| \(\hat{\rho}_{b,\mathrm{hi}}\) | Regression band, high | g/cm³ |
Single-value calculator
Behavior
The prediction is a plane in the predictors, so each curve on the plot is a straight line. With the default fit each 0.01 v/v of neutron porosity lowers the predicted density by 0.013 g/cm³, and the three gamma-ray values shift the line up as the gamma ray rises: at a neutron porosity of 0.2 v/v the predicted density is 2.319 g/cm³ at 30 gAPI, 2.411 at 70 and 2.526 at 120. The band at the default residual of 0.026 g/cm³ is ±0.051 g/cm³, so the density at 70 gAPI and 0.2 v/v is 2.411 with a band of 2.360 to 2.462 g/cm³. What the calculator cannot show is the main failure: if the neutron reading in the gap is itself damaged (it usually is), the neutron coefficient carries the damage straight into the prediction.
Parameter guidance
Predictors. Use curves that are not damaged by the same hole: gamma ray, deep resistivity and the sonic are the usual choices for density. The neutron and PE are affected by the same washout, and a model that uses them reproduces the damage; the worked example shows how much. Use the minimum number of predictors that give a good fit, and avoid highly collinear ones. Check the relation of each predictor to the target in a scatter plot first.
Training window. Fit on good samples from the same well, the same formation and nearby depth, not the whole well. A window of about 50 to 200 ft either side of the gap, excluding the nulled, padded samples, is a typical start. Extend the window if there are too few samples (hundreds of samples is comfortable for a few predictors) and shorten it if the relation visibly changes with depth.
Fit quality. Report R² and the residual standard deviation, and test on samples that were held out in blocks of depth, not randomly. The fit should leave a residual that is small against the range of the curve in the interval. A residual comparable to the signal means the repair adds little information.
Before relying on it. Compare with the real density wherever it is good, and compare the repaired intervals with nearby wells and with core. Keep the repair flag. The shared steps of flagging and padding are on Repair: Null-Out; a non-linear alternative is on Repair: Machine-Learning Infill.
Worked example
A synthetic 600 ft well in which shale volume and porosity drive the gamma ray, neutron, sonic and density, with a washout in which the density reads low and the neutron reads high. The density is nulled over the gap plus padding and repaired by regression on three sets of predictors, trained on a window about 90 ft either side. Because the data are synthetic, the true density in the gap is known and the repair can be scored.
rng = np.random.default_rng(12)
n = 1200 # 600 ft at 0.5 ft
smooth = lambda a, k: np.convolve(a, np.ones(k) / k, mode="same")
# Synthetic well: shale volume and porosity drive every curve
vsh = np.clip(0.5 + 3.0 * smooth(rng.normal(0, 1, n), 25), 0, 1)
phi = np.clip(0.28 - 0.18 * vsh + 0.02 * smooth(rng.normal(0, 1, n), 9) * 3, 0.02, 0.35)
gr = 20 + 100 * vsh + rng.normal(0, 6, n)
nphi = phi + 0.12 * vsh + rng.normal(0, 0.012, n)
dt = 56 + 120 * phi + 20 * vsh + rng.normal(0, 2.5, n)
rhob_true = 2.65 - 1.65 * phi + 0.02 * vsh + rng.normal(0, 0.012, n)
# A washout from sample 500 to 560: density reads low, neutron reads high
gap = np.arange(500, 560)
rhob = rhob_true.copy()
rhob[gap] -= np.linspace(0.15, 0.35, gap.size)
nphi_obs = nphi.copy()
nphi_obs[gap] += 0.06 # neutron is also affected
good = np.ones(n, bool)
good[485:575] = False # null-out with 15 samples of padding on each side
window = np.zeros(n, bool)
window[300:485] = True
window[575:760] = True # training window: about 90 ft either side of the gap
train = good & window
def fit(cols, mask):
A = np.column_stack([np.ones(mask.sum())] + [c[mask] for c in cols])
beta, *_ = np.linalg.lstsq(A, rhob[mask], rcond=None)
res = rhob[mask] - A @ beta
r2 = 1 - res.var() / rhob[mask].var()
return beta, r2, res.std(ddof=len(beta))
def predict(beta, cols, idx):
return beta[0] + sum(b * c[idx] for b, c in zip(beta[1:], cols))
models = {
"GR + NPHI": [gr, nphi_obs],
"GR + NPHI + DT": [gr, nphi_obs, dt],
"GR + DT (no neutron)": [gr, dt],
}
print(f"training samples: {train.sum()} (window {window.sum()}, nulled gap with padding {(~good).sum()})")
print(f"{'model':22s} {'R2':>6} {'RMSE':>7} {'gap bias':>9} {'gap RMSE':>9} coefficients (intercept, then predictors)")
for name, cols in models.items():
beta, r2, rmse = fit(cols, train)
pred = predict(beta, cols, gap)
err = pred - rhob_true[gap]
print(f"{name:22s} {r2:6.3f} {rmse:7.4f} {err.mean():9.4f} {np.sqrt(np.mean(err**2)):9.4f} " + " ".join(f"{b:.5g}" for b in beta))
# Same model with predictors that were not affected by the washout (neutron taken from the clean copy)
beta, r2, rmse = fit([gr, nphi], train)
err = predict(beta, [gr, nphi], gap) - rhob_true[gap]
print(f"\nGR + NPHI with an unaffected neutron: gap bias {err.mean():.4f}, gap RMSE {np.sqrt(np.mean(err**2)):.4f}")
print(f"For reference, RMSE of the damaged measurement itself in the gap: {np.sqrt(np.mean((rhob[gap] - rhob_true[gap])**2)):.4f}")
print(f"Approximate 95% band of the repair = +/- {1.96 * rmse:.3f} g/cm3")
Output
training samples: 370 (window 370, nulled gap with padding 90)
model R2 RMSE gap bias gap RMSE coefficients (intercept, then predictors)
GR + NPHI 0.953 0.0263 -0.0853 0.0892 2.5095 0.0022861 -1.3011
GR + NPHI + DT 0.955 0.0258 -0.0730 0.0780 2.6426 0.0023603 -1.1195 -0.002058
GR + DT (no neutron) 0.924 0.0334 -0.0101 0.0374 2.7514 0.0029382 -0.0068756
GR + NPHI with an unaffected neutron: gap bias -0.0072, gap RMSE 0.0271
For reference, RMSE of the damaged measurement itself in the gap: 0.2568
Approximate 95% band of the repair = +/- 0.052 g/cm3
Assumptions and limitations
- The relation between the predictors and the target in the gap is the same as in the training window. A change of lithology, fluid or tool in the gap breaks it.
- The relation is linear in the predictors. Density versus neutron and sonic is close to linear over a range of porosity but not across lithologies.
- The predictors are good in the gap. A predictor that is damaged by the same hole carries the damage into the repair.
- The training samples are representative of the gap, and independent of it. They are neighbours, so the training residual is optimistic about the error in the gap.
- The predicted values carry no new information about the formation. They are an interpolation among the other curves and so cannot be used as independent evidence in a later calculation that uses the same curves.
QC checks
- R² and residual standard deviation are reported, and the fit is checked on held-out depth blocks.
- The repaired curve joins the measured curve smoothly at the edges of the gap, and the junction is not a step (a step means a bias).
- The repaired curve stays inside the range of the training data; values outside it are extrapolations.
- Repaired intervals are marked by a flag curve and are listed in the project record.
- A different choice of predictors changes the repair by less than the band. If it does not, the repair depends on the model and is not stable.
- Where a coal, salt or other unusual rock is in the gap, the repair is not used.
Going Deeper
Regression repair has a long history in log editing and is a simple example of interpolation among correlated measurements. Its weakness is that the information in the repaired curve comes entirely from the predictors, so any later calculation that combines the repaired density with those same predictors (a density-neutron porosity, a mineral model) gains nothing from the repaired value. Where a damaged curve is repaired from a curve that was itself affected by the hole, the correlation is between two errors. The sensible use is to keep a continuous curve for visual and statistical work, and to mask or discount the repaired intervals in quantitative results.
References
- Asquith, G. and Krygowski, D., 2004. Basic Well Log Analysis, 2nd edition. AAPG Methods in Exploration Series 16, American Association of Petroleum Geologists, Tulsa, OK.
Python reference implementation
Python reference implementation
The Python reference implementation is available to registered users with a verified email address. Register or sign in to view it.