Combining TOC Estimates
On this page
Summary
Several independent estimates of Total organic carbon can be combined into one curve. The mean uses all of them, the median ignores a single outlier, and the minimum is conservative. Use a combination when estimates from different logs agree in trend but differ in detail, and when no one of them is known to be best.
Inputs and outputs
| Item | Units | |
|---|---|---|
| Input | TOC estimate 1 | wt% |
| Input | TOC estimate 2 | wt% |
| Input | TOC estimate 3 | wt% |
| Output | Mean TOC | wt% |
| Output | Median TOC | wt% |
| Output | Minimum TOC | wt% |
Equations
For \(N\) estimates \(TOC_1, \ldots, TOC_N\), the combined values are:
A weighted mean, \(\sum_i w_i\,TOC_i\) with \(\sum_i w_i = 1\), is the general form when the estimates are not equally reliable. The calculator below uses three estimates.
| Symbol | Variable | Units | Typical range |
|---|---|---|---|
| \(\mathrm{TOC}_1\) | TOC estimate 1 | wt% | 0 to 15 |
| \(\mathrm{TOC}_2\) | TOC estimate 2 | wt% | 0 to 15 |
| \(\mathrm{TOC}_3\) | TOC estimate 3 | wt% | 0 to 15 |
| \(\overline{\mathrm{TOC}}\) | Mean TOC | wt% | 0 to 15 |
| \(\mathrm{TOC}_{med}\) | Median TOC | wt% | 0 to 15 |
| \(\mathrm{TOC}_{min}\) | Minimum TOC | wt% | 0 to 15 |
Single-value calculator
Behavior
The plot holds two estimates at 3.2 and 3.8 wt% and sweeps the third. The mean follows the third estimate in a straight line, so one wild value pulls the result away. The median stays between the other two estimates once the third is the outlier, and the minimum stays at the lowest value.
Parameter guidance
The choice between the three is a choice about risk. The mean is appropriate when the estimates have independent, random errors. The median is the safe choice with three or more estimates and a possible outlier. The minimum is conservative: it is useful when each method can only overestimate, for example when the main errors are pyrite in density and hydrocarbons in resistivity. Before combining, calibrate each estimate against core individually, as covered on the TOC Analysis page. Estimates that share an input, such as two density-based ones, are not independent, and averaging them gives less than it seems to.
Worked example
Three estimates of 3.2, 3.8 and 7.5 wt%, where the third is an outlier:
import statistics
est = [3.2, 3.8, 7.5]
print(f"mean = {statistics.mean(est):.2f} wt%")
print(f"median = {statistics.median(est):.2f} wt%")
print(f"min = {min(est):.2f} wt%")
Output
mean = 4.83 wt%
median = 3.80 wt%
min = 3.20 wt%
Assumptions and limitations
- Each estimate has been calibrated against core, so a difference between them reflects error and not a different scale.
- The estimates have different error sources. Averaging two estimates that are driven by the same log reduces the error much less than expected.
- A single method is not already known to be best in the interval. If it is, use it.
QC checks
- The combined curve lies within the range of the inputs and is smoother than the worst of them.
- Where the estimates disagree by a large amount, find out why before trusting the combined result.
- The combined curve compares with core better than at least some of the individual estimates. If it does not, the combination is not helping.
Going Deeper
The alternative to a fixed average is to fit the combination to core, by regression of core TOC on the available logs or on the individual estimates. This learns the weights from data, at the cost of needing enough core and risking overfitting. Machine-learning versions of the same idea are common. Some workflows average a set of published density-based regressions, which smooths the differences between them but, as the estimates share an input, does not average out errors in the density log itself.
References
References will be added once verified.
Python reference implementation
Python reference implementation
The Python reference implementation is available to registered users with a verified email address. Register or sign in to view it.