CamPetro

Mnemonic Conventions and Curve Families

On this page

Summary

A Curve mnemonic names a curve only by convention, so the same measurement turns up under many names. This page shows how to define Curve family groups, build an Alias dictionary, normalise the names in a file and choose one curve per family when several qualify. Use it at the start of every project, before any curve is read by name.

Inputs and outputs

Item Notes
Input Curve mnemonics from the file As written, including suffixes such as :1 or _RAW
Input Unit string of each curve From the header, which may be missing or wrong
Input Sample statistics of each curve Fraction of valid samples, range, median
Input An alias dictionary Family, known names, expected units, priority
Output One chosen curve per family With the reason it was chosen
Output Ranked alternates Kept, not discarded
Output A list of unresolved curves For a person to look at

Equations

There is no governing equation. The procedure has six steps.

  1. Define the families you need (density, neutron porosity, compressional slowness, gamma ray, caliper, deep, medium and shallow resistivity, and so on), each with its expected units and a plausible value range.
  2. Normalise every mnemonic: lower case, strip repeat-section markers (:1) and duplicate counters (_1), turn separators into one form, and peel off a processing tag (_RAW, _EDIT) that says which version of the curve it is.
  3. Look the normalised name up in the alias dictionary. A name that matches nothing is reported, never guessed.
  4. Check the match: units must be in the family's list, and the values must fall in its plausible range. A curve with the right name and wrong units is converted later, not dropped, but it is flagged.
  5. When several curves match one family, rank them with a score. The one used in the example below is
\[ S = f_{\mathrm{valid}} + 0.1\,t - 0.5\,u \]

where \(f_{\mathrm{valid}}\) is the fraction of samples that are valid, \(t\) is \(+1\) for an edited or corrected version, \(-1\) for a raw one and \(0\) otherwise, and \(u\) is \(1\) when the units are not in the family's list and \(0\) otherwise. The weights are illustrative, not a standard.

  1. Keep the ranked alternates and record the choice, so that it can be reviewed and reversed.

Single-value calculator

No calculator: this method is a procedure, not a single equation.

Behavior

The run below resolves fifteen curves. Where several curves compete, coverage and the processing tag decide. RHOB and the repeat section DEN:1 tie at a score of 0.97 and the ordering between them is arbitrary, which is why a tie must be sent to a person. TNPH_EDIT (1.03) beats NPHI (0.95) and the raw TNPH_RAW (0.85). DTCO (0.99) beats DT (0.90), and ILD (0.98) beats an LLD that is valid over only 31% of the interval. TENS and XYZ1 do not match any family and are reported as unresolved. The scores are computed by the code, not tuned by hand.

Parameter guidance

Alias dictionary. Start from the mnemonics you meet in your own data, not from a published list, and grow it as new files arrive. Keep one row per family with a canonical name, the known mnemonics, accepted units and a plausible value range. Keep deep, medium and shallow resistivity as separate families, and keep induction and laterolog types separate, because they respond differently in the same rock.

Priority between duplicates. Use the score as a starting point and adjust the weights to your own rules. Typical rules: prefer an edited curve to a raw one, prefer more valid samples, prefer a curve whose units match the family, and prefer the higher-resolution curve unless it is noisy for the purpose. Do not rank by mnemonic alone. A tie is a reason to ask, not to pick.

Minimum coverage. A curve valid over only a small part of the interval is better treated as a patch for the main curve (see Curve Splicing and Merging) than as the chosen curve. Pick the threshold for the interval of interest, not for the whole file.

Vendor and tool specifics. The names of array channels and tool-specific variants are on Vendor and Tool-Specific Mnemonics. The choices shared across the stage are on the step page.

Worked example

Fifteen curves as they might arrive from a merged file, each with its unit string and the fraction of the interval over which it has valid data. The code is a complete, small resolver:

import re

# 1. A small alias dictionary: canonical family -> known mnemonics (lower case, no separators)
#    plus the units a valid curve of that family may carry.
FAMILIES = {
    "density":    {"names": {"rhob", "rhoz", "den", "dens", "denb", "bdens"}, "units": {"g/cc", "g/cm3", "kg/m3"}},
    "neutron":    {"names": {"nphi", "tnph", "nphil", "cnl", "npor", "nphils"}, "units": {"v/v", "pu", "%", "dec"}},
    "sonic_comp": {"names": {"dt", "dtc", "dtco", "ac", "dtcomp"}, "units": {"us/ft", "us/m", "usec/ft"}},
    "gr":         {"names": {"gr", "gra", "grc", "sgr", "ggr"}, "units": {"gapi", "api"}},
    "caliper":    {"names": {"cali", "cal", "hcal", "cal1", "dcal"}, "units": {"in", "mm", "inch"}},
    "res_deep":   {"names": {"rt", "ild", "lld", "rd", "resd", "rild", "at90"}, "units": {"ohm.m", "ohmm", "ohm-m"}},
}
# Tags appended by processing; they say which version of a curve this is, not what it measures.
TAG_RANK = {"raw": -1, "edit": 1, "edt": 1, "cor": 1, "fnl": 1}


def normalise(mnemonic):
    """Lower case, drop repeat-section markers (:1) and duplicate counters (_1), split off a processing tag."""
    m = mnemonic.strip().lower()
    m = re.sub(r":\d+$", "", m)
    m = re.sub(r"_\d+$", "", m)
    m = m.replace("-", "_").replace(" ", "_")
    tag = 0
    for t, rank in TAG_RANK.items():
        if m.endswith("_" + t):
            m, tag = m[: -len(t) - 1], rank
    return m.replace("_", ""), tag


def family_of(base):
    for fam, spec in FAMILIES.items():
        if base in spec["names"]:
            return fam
    return None


def resolve(curves):
    """curves: list of (mnemonic, unit, fraction_valid). Returns {family: ranked candidates}, unresolved."""
    found, unresolved = {}, []
    for mnem, unit, valid in curves:
        base, tag = normalise(mnem)
        fam = family_of(base)
        if fam is None:
            unresolved.append(mnem)
            continue
        unit_ok = unit.strip().lower() in FAMILIES[fam]["units"]
        score = valid + 0.1 * tag - (0.5 if not unit_ok else 0.0)
        found.setdefault(fam, []).append((score, mnem, unit, valid, unit_ok))
    for fam in found:
        found[fam].sort(reverse=True)
    return found, unresolved


if __name__ in ("__main__", "worked_example"):
    curves = [
        ("RHOB", "G/CC", 0.97), ("RHOZ", "G/CC", 0.64), ("DEN:1", "KG/M3", 0.97),
        ("NPHI", "V/V", 0.95), ("TNPH_RAW", "V/V", 0.95), ("TNPH_EDIT", "V/V", 0.93),
        ("DT", "US/FT", 0.90), ("DTCO", "US/FT", 0.99),
        ("GR", "GAPI", 1.00), ("SGR", "GAPI", 0.99),
        ("HCAL", "IN", 1.00), ("ILD", "OHMM", 0.98), ("LLD", "OHMM", 0.31),
        ("TENS", "LBF", 1.00), ("XYZ1", "", 0.50),
    ]
    found, unresolved = resolve(curves)
    for fam, cands in found.items():
        best = cands[0]
        print(f"{fam:11s} -> {best[1]:10s} (score {best[0]:.2f})")
        for score, mnem, unit, valid, unit_ok in cands[1:]:
            print(f"{'':11s}    alt {mnem:10s} score {score:.2f}, valid {valid:.0%}, units ok: {unit_ok}")
    print("unresolved:", ", ".join(unresolved))

Output

density     -> RHOB       (score 0.97)
               alt DEN:1      score 0.97, valid 97%, units ok: True
               alt RHOZ       score 0.64, valid 64%, units ok: True
neutron     -> TNPH_EDIT  (score 1.03)
               alt NPHI       score 0.95, valid 95%, units ok: True
               alt TNPH_RAW   score 0.85, valid 95%, units ok: True
sonic_comp  -> DTCO       (score 0.99)
               alt DT         score 0.90, valid 90%, units ok: True
gr          -> GR         (score 1.00)
               alt SGR        score 0.99, valid 99%, units ok: True
caliper     -> HCAL       (score 1.00)
res_deep    -> ILD        (score 0.98)
               alt LLD        score 0.31, valid 31%, units ok: True
unresolved: TENS, XYZ1

Assumptions and limitations

  • The name is a usable clue. A mnemonic may be reused by a vendor for a different quantity, so a name match is a proposal that must be confirmed by units and values.
  • The unit string in the header is correct. Many are blank, wrong or inherited from a template.
  • The scoring weights are acceptable. They are illustrative and will not suit every project.
  • The families are complete for the project. A new tool type needs a new family or a new alias.

QC checks

  • Every curve the workflow will use has exactly one chosen source, and the choice is listed with its alternates.
  • The list of unresolved curves was read by a person. Some of them (tension, speed, temperature) are harmless, others are the curve you needed.
  • The values of the chosen curve fall in the plausible range of its family. A density of 0.0 to 1.0 or a gamma ray of 10 000 is a wrong match or a wrong unit.
  • The chosen curves are mutually consistent: the neutron and density come from the same pass and the same depth reference.
  • The same file, loaded twice, gives the same choices.

Going Deeper

Mnemonic standards exist, but none is universal. The LAS file format carries a mnemonic, a unit, a value and a description for each curve, and the description is often more informative than the mnemonic. Vendors publish their own mnemonic catalogues, and these change with each tool generation, so a dictionary built from catalogues is always incomplete for older and merged data. A fully automatic resolver can therefore never be trusted without a review. The practical compromise is a dictionary that is conservative about what it matches, reports what it cannot match, and records every decision so that it can be audited.

References

  1. Canadian Well Logging Society, 1992 (revised 1999). LAS Version 2.0: A Digital Standard for Logs. Update February 1992 (CWLS Log ASCII Standard).
  2. Rider, M. and Kennedy, M., 2011. The Geological Interpretation of Well Logs, 3rd edition. Rider-French Consulting Ltd, Sutherland, UK.

Python reference implementation

Python reference implementation

The Python reference implementation is available to registered users with a verified email address. Register or sign in to view it.