Large Labor Model
WHEN WILL AI REACH YOUR JOB?

Methodology

Version 5.0 — Released April 19, 2026 · Historical backfill and late fixes applied April 20, 2026 · Last revised August 17, 2026 (changelog) · capability snapshot: April 18, 2026 · timeline current through August 2026.

This document describes how the Large Labor Model v5 dataset was produced. It is written so that a reader with no prior knowledge of the project can understand how every number was derived, what assumptions were made, and where the model is weakest. Nothing in the dataset is inherited from prior versions unexamined; v5 is a ground-up rebuild on a reviewed 13-phase pipeline.

Companion documents, all shipped in the repository: the full changelog, the methodology reference markdown, and the editorial process record documenting the emergency interventions during Phase 11.

1. What this measures

The Large Labor Model maps two separate things over time:

Labor shares. How many people work in each category of human activity, globally, from 1800 to 2041. Historical data from Bairoch, Maddison, Mitchell, GGDC 10-sector, and ILOSTAT. Forward projections from the displacement model.

Replaceability. When AI or robotics becomes technically capable of performing the core tasks of each occupation, scored 0–100. Computed from the six-vector task decomposition. Covers 1970 → 2041 for modern territories.

These are separate measurements. Replaceability tracks the technical frontier. Labor shares track the economy. The gap between the two — between what AI can do and what has actually changed in the labor market — is not a flaw in the model. It is one of the most important things the model shows.

1.1 Replaceability ≠ Replacement

The load-bearing editorial distinction of v5, refined across the build and corrected in Phase 11 after a review-caught conflation:

Replaceability = tech capability × economic viability × commercial availability. A high score means the technology exists and works on the occupation's tasks, the unit economics are plausible at realistic wage points, and the product is available to buy today. It does not bake in regulatory or preference barriers.

Replacement = adoption lag × regulatory friction × human-preference inertia × capital cycles. This is the dynamic question of how fast replaceability translates into actual workforce displacement. It lives in the Phase 6 replacement formula, not in the replaceability score.

The canonical example: a robo-barista. Capability exists, economics can work, products are on the market — replaceability is high. But customers still prefer human baristas, many jurisdictions regulate food service, and adoption moves slowly. So replacement lags dramatically. The clean framing forces these two numbers to be read separately.

The same logic applies retrospectively. In 2010, POS + self-checkout technology made a large share of cashier tasks replaceable. Actual cashier displacement by 2014 was much smaller. That gap is barriers, not capability.

2. Territories and occupations

Human work is organized into 15 territories: 13 modern territories visible from the late 20th century onward, plus 2 historical aggregate territories (Industry and Services) and Land & Sea (which spans both eras). Each modern territory maps to one or more sections of the International Standard Industrial Classification (ISIC Rev. 4). Territory names are editorial — legible to a general audience — while the ISIC mapping keeps the underlying definitions standard and internationally comparable.

Assignment principle. Occupations are assigned by the nature of the work, not by the industry sector where the work happens. A barber's work is personal care, so barbers belong to Care & Health regardless of whether they work in a salon or a hotel. A car mechanic's work is repair, so mechanics belong to Maintaining & Fixing regardless of where they work. This resolves the ambiguity between industry and occupation classifications.

Functional managers. Managers whose role is defined by a specific function (Finance Managers, IT Managers, Marketing Managers) are assigned to the territory of their function — Money & Data, Money & Data, Making Meaning — rather than to Thinking & Leading. Thinking & Leading is reserved for organizational generalists (Chief Executives, Management Consultants), pure knowledge workers (Mathematicians, Sociologists), and people functions (HR Specialists).

TerritoryWhat it coversISIC Sections
Land & SeaAgriculture, forestry, fishing, mining, quarryingA, B
Making ThingsManufacturingC
Building ThingsConstructionF
Moving ThingsTransport, logistics, warehousingH
Buying & SellingRetail and wholesale tradeG
Money & DataFinance, insurance, real estate, informationJ, K, L
Care & HealthHuman health and social workQ
Learning & TeachingEducationP
Making MeaningArts, media, creative industriesR
Governing & ProtectingPublic administration, defenseO
Feeding & HostingAccommodation, food serviceI
Maintaining & FixingUtilities, repair, personal servicesD, E, S
Thinking & LeadingProfessional, scientific, technical, managementM, N

2.1 480 occupations

v5 covers 480 occupations, up from 388 in v3.2. Each is anchored to an ISO standard ISCO-08 4-digit unit group code where one exists; 10 specialized splits use documented parent + suffix conventions (e.g., 2512-ml for ML engineers as a split from 2512 Software Developers). Every occupation carries a common-titles list drawn from O*NET, BLS, and live LinkedIn 2025–2026 job-posting scrapes; a display name; a territory assignment; and a task decomposition.

The Phase 2 coverage expansion added missing ISCO-08 unit groups and specialized splits where a single code conflates genuinely distinct real-world roles (ML engineers vs. general software developers, tax accountants vs. general accountants, network operations technicians vs. general IT support). The split convention preserves traceability: the parent ISCO code is always preserved, so specialized splits can be rolled up into their parent without data loss.

2.2 Task decomposition

Every occupation is decomposed into 3–14 tasks (median ~10), each carrying:

Phase 4 produced the task set by merging three independent retrieval sources — O*NET task statements via SOC → ISCO crosswalk (15,575 raw tasks), BLS Generalized Work Activities + OOH (11,234 raw tasks), and a context-free reconciler that applied canonical vector definitions without occupation priors. Three-way reconciliation achieved 84.2% agreement; split cases were defaulted to the adversarial pass and flagged. Phase 11 ran three adversarial audit rounds that removed 825 scrape artifacts in aggregate and reclassified hundreds of tasks whose vector assignments were wrong in specific patterns (physical work tagged cognitive; field service tagged industrial; performance work tagged generative).

Total inline task decompositions: 4,811 tasks (4,818 before the June 8 duplicate-cleanup). Each carries a source attribution (O*NET, BLS, or dual-sourced) and participates in the replaceability computation described in Section 4.

3. The six capability vectors

Every AI/robotics capability is projected onto one of six vectors. These were chosen because they produce distinct deployment curves — cognitive-routine work scales with software ubiquity, physical automation scales with industrial capital, driving scales with sensor + map + insurance regimes, unstructured physical scales with humanoid hardware. Collapsing them into a single "AI capability" number obscures the divergent deployment stories.

VectorWhat it measures2026 value
C_R — Routine CognitiveCodifiable procedural cognitive work: data entry, filing, routine coding, form processing, ledger keeping, rule-based medical coding0.76
C_G — Generative CognitiveOpen-ended judgment and creation: novel analysis, synthesis, strategy, original writing, novel diagnosis, senior legal/medical reasoning0.57
P_A — Physical AutomationHighly repeatable physical motion in controlled, engineered environments: factory assembly, CNC, plant-line machinery0.75
Phi_S — Selective PhysicalSensor-based structured-environment physical: route-based driving (Waymo), warehouse picking, mail delivery0.46
Phi_U — Unstructured PhysicalUnpredictable real-world physical: surgery, in-home plumbing, home care, live embodied performance, variable-terrain farm work0.15
S_E — System EngineeringMulti-agent/multi-stakeholder orchestration: enterprise architecture, platform engineering, air traffic control, emergency coordination0.35

The 2026 values are the Phase 11 post-recalibration set. Phase 5's original blind calibration produced higher benchmark-anchored values (C_R 0.87, C_G 0.62, P_A 0.78, Phi_S 0.60, Phi_U 0.15, S_E 0.42). Phase 11 ran two independent recalibrators (Claude Opus 4.7 and GPT-5.4) under the same brief: re-interpret each benchmark citation as a production-reliability number, discounting for remote-operator backstops, confident-failure rates, narrow operational-design domains, and agentic brittleness. Both instances converged within ±0.03 on every vector. The shipping set is the blended midpoint.

The largest correction is Phi_S (−0.14). This is the vector where benchmark and demonstration evidence most overstates production capability under honest-deployment accounting. Waymo's weekly paid rides — roughly 250K at the April snapshot, ~500K by August 2026 — are real deployment, but they run narrow operational-design domains and depend on remote-operator backstops for edge cases. Phi_U is defended unchanged at 0.15 because it was already at the conservative floor in Phase 5 — no 2026 deployment evidence moves it.

Vector ordering (C_R > P_A > C_G > Phi_S > S_E > Phi_U) is preserved across the recalibration. Ordering is load-bearing: if it inverted, something would be wrong with the deployment evidence or the vector definitions.

4. How replaceability is scored

For each occupation and each year, replaceability is computed by comparing each task's capability vector value (at that year) against the task's difficulty threshold. Tasks that cross the threshold contribute their time weight to the score.

replaceability(occupation, year) = 100 · Σtasks time_weight(task) · crossover(capability[task.vector][year], task.difficulty) ÷ Σtasks time_weight(task) where crossover(cap, diff) = 1 / (1 + exp(-8 · (cap − diff)))

The logistic crossover (k=8) gives partial credit as capability approaches the task difficulty, rather than a strict all-or-nothing threshold. This is the right shape both below the threshold (early partial automation — ATMs did a fractional share of teller work well) and above it (no single deployed system reaches 100% reliability). The same formula is applied across the full 1970 → 2041 window for consistency.

At capability = difficulty the crossover value is 0.5 (half credit). At a +0.25 gap above the threshold it reaches 0.88; at a −0.25 gap below it drops to 0.12. The k=8 slope was chosen to match honest deployment behaviour — systems rarely snap from "doesn't work" to "fully works" at a single capability increment, but they also don't drift to full capability over arbitrary gradient scales.

k=8 is a convention, not an empirical constant. The data does not uniquely identify a slope; k=8 is a design choice in the range most consistent with observed labor automation patterns. k=4 would compress scores toward the middle and damp noise from small capability or difficulty edits; k=12 would behave closer to a hard threshold and amplify task-decomposition noise. k=8 sits in between. Occupations with many tasks whose difficulty is near the capability value are sensitive to this choice — if a future revision changes k, some borderline occupations will shift several points while occupations with task difficulties far from capability will barely move.

Territory aggregation. Territory-level replaceability is the employment-weighted mean of in-territory occupation replaceabilities, using current (2026) in-territory employment weights. Weights come from ILOSTAT 2025, BLS OES, and mapped via SOC → ISCO. For historical years, this produces "what share of today's labor composition, within each territory, would have been technically replaceable at that year's capability values" — a deliberate editorial choice documented at metadata.replaceability_computation.

Every territory × year replaceability score in the dataset is reproducible from: (a) the Phase 4 task decomposition, (b) the reconciled 1970 → 2041 capability trajectory (Phase 7 forward + Phase 12 backward, post-shift), (c) the logistic crossover formula above, and (d) the employment weights. The metadata.replaceability_computation block records the formula and the inputs so the chain is walkable end-to-end.

4.1 Global replaceability

The global replaceability metric shown in the visualization header is the global-workforce-weighted mean of territory replaceabilities — 33.7 for 2026. The unweighted mean across the 480 occupations is higher, r26 = 41.2, because employment concentrates in less-replaceable physical work. Under the logistic crossover (k=8) the ceiling is asymptotic — no occupation reaches exactly 100. The most-replaceable score is around 96 (clerical roles such as data-entry clerks where every task's difficulty threshold sits well below C_R = 0.76). At the other end, 4 occupations score r26 ≤ 5 and 22 score below 10 — humanoid-dependent physical roles (surgeons, plumbers, child care workers) where Phi_U = 0.15 is well below almost every task threshold.

5. The replacement model

Replaceability is converted into actual workforce displacement through a calibrated formula that captures three distinct frictions: adoption lag, barrier-type-specific drag, and era-stratified already-absorbed displacement.

displacement(t) = min( (conversion_rate × (replaceability_t − replaceability_2026) × lag(t)) × barrier_multiplier(primary_barrier) + already_absorbed(occupation, era), replaceability_t)
ParameterValueNotes
conversion_rate0.30Fraction of replaceability-capability gain that translates into workforce reduction at steady state. Fit against 7 historical automation cases (telephone operators 1920–2023, bank tellers, travel agents, cashiers, translators, customer service reps, bookkeeping clerks).
lag schedule2026: 7%
2030: 18%
2035: 30%
2040: 42%
2041: 44%
Piecewise, interpolated between anchors. Slower than v3.2 (5 / 28 / 48 / 58 / 60). Historical automation cases averaged 15+ years from capability emergence to majority displacement; v5 honors that.
barrier multipliersNONE: 1.00
ECONOMIC: 0.95
HUMANOID_DEPENDENT: 0.75
HUMAN_PREFERENCE: 0.50
REGULATORY: 0.35
Strengthened from Phase 6's original (0.70 and 0.50) during Phase 11 to preserve fit against the Anthropic Economic Index cross-section after r26 was rescored to pure task-math. Cross-section R² rises from 0.61 plain → 0.85+ with barrier stratification.
already_absorbedera-stratifiedPer-occupation floor representing displacement that already happened before 2026 (e.g., telephone operators, typists). Prevents the model from "re-displacing" already-absorbed cases. Era assigned by a barrier-type heuristic (NONE → post_2022; REGULATORY/HUMAN_PREFERENCE → y2015_2022; ECONOMIC/HUMANOID_DEPENDENT → pre_2015). The heuristic is a rough prior: Accountants (REGULATORY → y2015_2022) likely understates already-absorbed because ERP/OCR/tax-software automation matured pre-2015; Cooks (HUMANOID_DEPENDENT → pre_2015) likely overstates it because the binding capability is embodied dexterity, not pre-2015 software. A per-occupation era override is a v5.1 enhancement.

Parameters were fit against the 7 historical automation cases and the 28-occupation cross-section from the Anthropic Economic Index (Massenkoff & McCrory, March 2026). Historical-case R² under the fitted formula is 0.85+; plain (no-barrier, uniform-absorbed) R² is 0.61.

Mean 2041-moderate displacement across 480 occupations is 24% (24.2), range 10–49 — carried per-occupation as displacement_2041_moderate. This is the honest interpretive anchor for 2041: replaceability rises across most occupations but displacement carries the variance.

6. Capability trajectories 1970 → 2041

Why 2041. The horizon is now + 15 (2026 + 15 = 2041). Fifteen years is where, for technology moving at this cadence, model projection shades into speculation; the model refuses to score past it. The horizon advances with the annual recalibration.

v5 carries a full 1970 → 2041 trajectory for every capability vector. The forward portion (2026 → 2041) came from Phase 7 — two parallel forecasters (Claude + GPT) researching published capability curves, humanoid shipment forecasts, benchmark trajectories, and deployment pilots; reconciled under mid = min(A_mid, B_mid) so no single forecaster's aggressiveness pulled the central case. The backward portion (1970 → 2014) came from Phase 12 — same reconciliation discipline, same two-instance structure, parallel honest-deployment research over assembly-line through ChatGPT milestones.

YearC_RC_GP_APhi_SPhi_US_E
19700.070.0050.100.0050.020.005
19850.160.0090.250.020.020.011
20000.360.020.420.050.030.05
20100.470.040.510.070.040.09
20200.590.080.650.210.060.18
20250.740.560.740.440.140.33
20260.760.570.750.460.150.35
20300.830.740.820.620.320.60
20350.860.820.890.720.540.73
20410.870.860.920.760.660.78

The shape is deliberate:

6.1 Phase 11 parallel shift

After the Phase 11 recalibration lowered 2026 anchors, the Phase 7 forward trajectory was parallel-shifted so that the 2026 anchor equals the recalibrated value while preserving the trajectory growth shape. Pre-shift Phase 7 mid values are retained in the trajectory file under *_phase7_original fields for traceability.

6.2 Phase 12 historical backfill — counterfactual back-cast

Read the historical series as a counterfactual, not a literal historical observation. The honest framing: "If today's task decompositions and today's territory employment mix had faced each historical year's capability levels, what would the implied replaceability have been?" That is coherent; it is not a claim about what labor markets literally looked like in 1970 or 1990. Task bundles were redefined by digitization. Occupational composition shifted. Some of today's occupations barely existed decades ago. The back-cast is a thought experiment that anchors the capability trajectory in real historical deployment evidence; it is not a reconstruction of historical labor.

Added 2026-04-20 after a spot-check surfaced that the v5.0 release was displaying 0% replaceability for every territory before 2015. Two instances (GPT-5.4 + Claude Opus 4.7) independently researched 1970–2025 with 15–25 URL-cited milestone anchors per vector: first commercial ATMs (1967–1970), CNC + microprocessor commercialization (1972–1975), VisiCalc / Lotus 1-2-3 (1979–1983), ERP precursors (SAP R/3 1992), IFR industrial-robot data series (1993), commercial web (1995), Deep Blue (1997), Google Translate (2006), Watson Jeopardy (2011), Kiva + Amazon Robotics acquisition (2012), AlexNet (2012), Kubernetes + RPA (2014), AlphaGo (2016), Transformer paper (2017), Waymo commercial (Dec 2018), GPT-2/3/ChatGPT/GPT-4 (2019–2023), Amazon Sequoia/Proteus (2023), humanoid pilots (2024). Both instances produced 56-year × 6-vector × 3-band trajectories; reconciler applied mid = min(A_mid, B_mid), pre-2015 cap at 0.85, and monotonicity enforcement. The reconciled trajectory was applied via logistic crossover to produce 585 historical territory-replaceability records (13 territories × 45 years).

6.3 The data/projection seam — 2025 (revised August 15, 2026)

The trajectory is built from two independently constructed regimes — the Phase 12 back-cast (≤ 2025) and the Phase 7/11 forward reconciliation (≥ 2026) — meeting at the calibrated 2026 anchor. As originally shipped, the seam was constructed naively: each band's historical series reached its 2026 anchor early and then repeated it. In the mid and low bands, the 2025 value duplicated the 2026 anchor for all six vectors; in the high band, C_R sat flat from 2024 and P_A from 2023. Every rendered curve therefore showed a one-to-three-year capability plateau immediately before the present — a counterfactual implication (that capability froze through 2025) contradicted by the model's own event timeline.

Correction (August 15, 2026). Each flat run was re-laid as a strictly increasing approach to the anchor. Single-year runs: v₂₀₂₅ = max(midpoint(v₂₀₂₄, v₂₀₂₆), v₂₀₂₆ − (v₂₀₂₇ − v₂₀₂₆)), clamped strictly inside (v₂₀₂₄, v₂₀₂₆) — the curve arrives at the anchor with approximately the slope it leaves with. Multi-year runs (high-band C_R and P_A, total moves ≤ 0.01): linear fill with strict increase enforced after rounding. Worked example, C_G mid: 2024 = 0.430 → 2025 = 0.520 → 2026 = 0.570, matching the forward first step of +0.050.

Scope. Eighteen series (6 vectors × 3 bands) were touched only within their flat runs. The 13 territory-level 2025 replaceability records were recomputed from the corrected trajectory; all moved down, by 1.0 (Maintaining & Fixing) to 7.2 points (Thinking & Leading). Invariants, verified by assertion: every year at or before a run's base is unchanged; every value from 2026 on is unchanged; every published 2026 figure is unchanged; all 18 series are strictly increasing through the seam (2014 → 2028).

Epistemic status of 2025, as first corrected. The corrected 2025 values were initially interpolated approach values between the last milestone-anchored historical year and the calibrated 2026 anchor — not independently researched anchors. The visible deceleration into 2026 reflects the honest-deployment discount applied to the anchor (benchmarks ran ahead of production reliability), not a claim that progress slowed. The correction is reproducible: v5-build/phase12/reconciler/seam_fix_aug2026.py and seam_fix_downstream.py, with machine-readable records in metadata.seam_fix_2026_08 and the Phase 12 trajectory file's own metadata.

Anchored (Phase 12b, later the same day). The interpolated 2025 values were then replaced by a three-instance research anchoring: Instance A (Claude family, live web retrieval, 44 URL-cited milestones), Instance B (GPT family, internal-knowledge sourcing, disclosed as such), and Instance C (GPT-5.6 Sol, run independently by the project owner under a de-biased protocol — ordinal judgment committed before seeing the numeric frame, and no tempered-tiebreak instruction). Reconciliation: per-band minimum across the three. The de-biased instance landed lowest on the cognitive vectors — evidence for the tempered read being evidence-driven rather than instruction-driven — and highest on the physical vectors, where narrow brackets bound the disagreement to ≤0.014. No instance reported bracket stress. Worked example, C_G mid: 2024 = 0.430 → 2025 = 0.487 → 2026 = 0.570. The 13 territory-level 2025 records were recomputed (−0.4 to −4.3 points versus the interpolation). Instance files, evidence, and reconciliation scripts: v5-build/phase12b_2025_anchor/; machine-readable records in metadata.phase12b_2025_anchor.

7. Three scenarios

Every occupation and every year carries low / mid / high bands for replaceability and displacement. The mid (moderate) scenario is the editorially-reviewed default — displayed on the canvas, cited in headlines.

2041 mid anchors (post-shift): C_R 0.87, C_G 0.86, P_A 0.92, Phi_S 0.76, Phi_U 0.66, S_E 0.78. 2041 high anchors: C_R 0.91, C_G 0.95, P_A 1.00, Phi_S 0.88, Phi_U 0.995, S_E 0.93.

Important caveat. v5.0 ships the moderate scenario as the editorially-reviewed display. Low and high band values are preserved per occupation as structural metadata, but they are produced deterministically from Phase 7's capability-trajectory low/high curves applied against Phase 4 task difficulties — they were not reviewed occupation-by-occupation. This produces editorially-implausible shapes in thin-margin cases (e.g., preschool educators' r30-high jumps from 8 to 100 when Phi_U high=0.67 crosses Phi_U task thresholds of 0.55-0.60). Re-examination of the low/high bands is a v5.1 deliverable. For now, reader attention should stay on the moderate band. metadata.scenarios_shipped encodes this disclosure.

8. The production pipeline

v5.0 was built on a reviewer-directed 13-phase pipeline. Each phase has a canonical output, a reviewer-signed checkpoint, and archived inputs/outputs. The build uses adversarial multi-instance separation: builders never validate their own output; separate instances do auditing, reconciliation, and documentation.

PhaseWhat it produces
0Schema lock + spec (v5_schema.json, Draft 2020-12)
1ISCO-08 canonical reference (436 unit groups + 10 approved custom codes)
2Occupation inventory + territory mapping (480 occupations locked)
3Labor data rebuild with occupation-up methodology (635 records, 1800–2025)
4Task decomposition (3-way reconciled; 4,818 tasks; 4,811 after the June 8 duplicate-cleanup)
5Replaceability 2026 scoring (blind calibrator + 3 scorers + reconciler)
6Replacement formula calibration (7 historical cases + 28-occupation cross-section)
7Capability trajectories 2026 → 2041 (2 forecasters + reconciler, three-scenario)
8Barrier notes + common titles (3-source merge + targeted anti-template rerun)
9Timeline events audit (researcher + verifier)
10Adversarial validation (isolated instance, 25 required checks)
11Editorial sign-off + emergency interventions (see editorial process record)
12Historical capability-vector backfill 1970 → 2014

8.1 Phase 11 emergency interventions

Phase 11 was supposed to be editorial sign-off + release. Instead it surfaced multiple deep issues that required intervention before ship. A full post-mortem is in the editorial process record. Headline interventions:

These interventions matter because scrape artifacts, vector mis-calibration, and the benchmark-vs-deployment gap are systemic issues that could have shipped undetected. The 2-audit + dual-recalibration process at Phase 11 caught them. This is why the project uses adversarial multi-instance separation.

9. Data sources

Primary sources backing v5:

A full source registry is inlined in the dataset at sources[]. Every labor record, every occupation, and every task carries source_ids[] pointing to registry entries with id, type, title, reliability, authority, and citation. Total source entries: 6,704.

10. Known limitations

This is a structured estimate, not a forecast. The honest statement of what the model cannot do:

  1. Point estimates over heterogeneous distributions. C_G spans first-draft copywriting to frontier diagnosis. Phi_S spans robotaxis to rural mail to heavy equipment. S_E spans coding agents to multi-stakeholder incident command. A single 2026 value per vector averages over this; the bounds carry some of the uncertainty but a vector split is a likely v6 improvement.
  2. Benchmark-vs-deployment gap. Benchmarks measure frontier capability on clean tasks; deployment measures production reliability on messy real labor. Phase 11's recalibration was an explicit attempt to close this gap via an honest-deployment lens, but production reliability data is still developing and values may shift in v5.1 / v6.
  3. High 2040-mid shape. Under the moderate scenario the unweighted occupation mean of technical replaceability at 2040 is 84.6 (employment-weighted, recomputed live from the trajectories: 84.1) and 25% of occupations score ≥90. The logistic crossover ceiling means no occupation reaches exactly 100 — the most-replaceable score is around 98. Most occupations sit in the 80–95 band by 2040, so the model still pushes interpretive weight onto the displacement column where the variance lives. Use displacement (mean 24%) to understand workforce implications, not replaceability alone.
  4. Phase 6 barrier-model form. Multiplicative barrier multipliers work across most of the distribution but produce counterintuitive results at high r26 × REGULATORY (lawyers at r41 = 100). A floor-adjusted barrier model for regulatory frictions is a likely v6 change.
  5. Residual scrape artifacts below the audit bar. Four Phase 11 audit rounds removed 825 artifacts. The audit bar was 2-model agreement; below-threshold contamination likely remains. A v6 Phase 4 re-run with tightened anti-scrape guardrails is recommended.
  6. Task-classification ceiling. Phase 4's three-way reconciliation achieved 84.2% agreement; the 15.8% split cases were defaulted to the Codex adversarial pass and flagged.
  7. Global-weighted P_A is lower than the frontier. Industrial-robot density in Korea is 7× China's and 5× Western Europe's. P_A at 0.75 averages across the global workforce. A single global value trades off against locale-specific precision.
  8. HUMANOID_DEPENDENT capability gap is uncertain. The 2041 Phi_U mid at 0.66 assumes humanoid hardware matures from 0.15 in 2026. Low-band (0.29) is the conservative read; high-band (0.995) is the aggressive read. The central value encodes substantial forecaster uncertainty.
  9. Editorial overrides pending. Two evidence-backed concerns (ISCO 5245 Service Station Staff; REGULATORY-barrier occupations with r41 ≥ 95) are documented but not applied to v5.0. v5.1 / post-launch.
  10. Historical reconstructions are coarse. 1800–1869 relies on Bairoch/Maddison macro aggregates; 1870–1990 on GGDC 10-sector. Only 1991+ is occupation-level ILOSTAT. Pre-1991 territory totals are honest but not precise.
  11. Low/high scenario bands are structural, not editorially reviewed. See Section 7 caveat.
  12. Historical backfill uses current task decompositions and current within-territory employment weights. Phase 12 produces honest historical capability values, then applies them against today's occupation task lists weighted by today's in-territory employment shares. This gives "what share of today's labor would have been technically replaceable at that year's capability" — not "what share of 1985's actual labor force was replaceable by 1985 technology." Reconstructing historical occupation composition at the 481-occupation granularity is deferred.
  13. Pre-1970 has no territory replaceability series. Capability evidence pre-1970 is too thin to support per-territory numbers with the same discipline as 1970+. Pre-1970 shows labor composition via historical_occupations macro aggregates (land_sea / macro_industry / macro_services) but no territory-level replaceability scores.

11. Changelog

August 17, 2026 — 2041 checkpoint completion

The replacement formula has always specified a 2041 lag entry (lag_schedule["2041"] = 0.44, §5) and this page anchors displacement interpretation at 2041 — but the per-occupation emission loop stopped at 2040: the dataset carried displacement_2040_moderate as its last displacement field while territory aggregates already ran to 2041. Completed: replaceability_2041_moderate/low/high and displacement_2041_moderate added for all 480 occupations via the documented pipeline (logistic k=8 task-math over the post-shift forward trajectory; Phase 6 formula). The extension script first reproduced every stored 2040 field exactly (480/480 within rounding) before emitting 2041. Results: mean 2041 displacement 24.2 (range 9.6–48.7), employment-weighted mean 2041 replaceability 84.2, zero 2040→2041 monotonicity violations; §5's stated range updates 9–47 → 10–49. Script: v5-build/assembler/add_2041_checkpoint.py.

August 15, 2026 — seam correction, currency patch, count fixes

Data/projection seam corrected (§6.3): flat runs where each band's history reached the 2026 anchor early were re-laid as strictly increasing approaches; the 13 territory-level 2025 records were recomputed (−1.0 to −7.2 points). No 2026 anchor, forward value, or published 2026 figure changed. Timeline currency patch: technology_events grew 77 → 94 (the May–August 2026 frontier wave, embodiment and regulatory milestones the timeline lacked, and two 2026 mathematics results — the Jacobian-conjecture counterexample and an amateur's GPT-5.4-assisted proof of a 60-year-old problem), each with a source entry; no occupation scores or capability values changed. The timeline display remains one dot per year, with one intentional second slot for 2026. Count fixes: metadata.source_count and occupation_count corrected to actual array lengths (423 → 6,704 entries; 481 → 480). This page: §4.1 previously conflated the workforce-weighted mean (33.7 for 2026) with the unweighted occupation mean (41.2); corrected. 2025 anchored (Phase 12b): later the same day, the interpolated 2025 values were replaced by a three-instance research anchoring (§6.3) — two model families, one instance under a de-biased protocol, per-band minimum — and the territory 2025 records were recomputed (−0.4 to −4.3). A standing, dated evidence ledger for future recalibrations was added at docs/evidence_ledger_2026H2.md.

June 8, 2026 — display-data polish (Phases 16–17)

Concise display labels (task_short) added for all 4,811 tasks (Phase 16). Mirror-surfaced data polish (Phase 17): four occupation score corrections, duplicate-task cleanup on ten occupations, 22 territory-year points adjusted by ±0.1. No change to the model's method.

v5.0 (April 2026)

Ground-up rebuild on a 13-phase pipeline. Every scoring column was recomputed; v3.2 served only as post-hoc sanity check and was never staged as input to any v5 phase.

April 21, 2026 — post-launch editorial polish

Two parallel review instances (Phase 14 legibility sweep and Phase 15 common-titles enrichment) surfaced 45 legibility findings and 11 common-titles enrichment proposals. After reviewer sign-off, applied: 33 display renames covering truncated labels (e.g., "Traditional & Complementary" → "Traditional & Complementary Medicine"; "Coding" → "Coding & Proofing Clerks"; "Aircraft Engine Mechanics" on ISCO 7234 → "Bicycle Repairers" matching the official ISCO scope), 11 common_titles updates filling search gaps for crypto/blockchain (homes at 3311, 2512, 2413), yoga / Pilates / wellness coaching (3423), modern AI roles (2512-ml), social-media managers (2431-social), drone operators (3153-uas), solar PV installers (7411-solar), and life / executive coaches (2635). Three previously-broken common_titles arrays (2431-social, 3153-uas, 2512-ml) had been carrying contamination from unrelated occupations and are now correct.

Two larger editorial moves applied in the same pass. ISCO 1219 ("Hospitality Directors") was renamed to "Admin Services Directors" and reassigned from feeding_hosting to thinking_leading — ISCO 1219 official scope is corporate admin/services managers (NEC), not hospitality industry executives, who already have separate homes at ISCO 1411 and 1412. Within-territory employment weights for both feeding_hosting and thinking_leading were renormalized using global_estimated_employment as ground truth, so each territory's weights still sum to 100% post-move. ISCO 3422-personal (an off-by-one specialized split for personal trainers) was dropped — personal training officially sits in ISCO 3423, which is now the canonical home for personal trainers, yoga instructors, Pilates instructors, group fitness, and wellness coaches. Occupation count: 481 → 480.

The pipeline was re-run end-to-end against the modified inputs: per-occupation replaceability at 2026/2030/2035/2040 (mid/low/high), displacement at 2026/2030/2035/2040 (Phase 6 formula), and territory aggregates 1970–2041 with dispersion stats at anchor years. Phase 10 validator: PASS. Two items were deferred to v5.1: thin-titles enrichment for ISCO 5211/5212/9520 (which lost contamination but need web-verified replacements), and a proposed specialized split for AML/KYC compliance officers. Both recorded at metadata.deferred_to_v5_1.

April 21, 2026 — formula coherence fix + Phase 13 verification

A Codex GPT-5.4 instance (Phase 13) independently verified the April 21 recompute. Reproduced every r / displacement / territory value at max_abs_diff = 0.0 across 480 × 7 replaceability fields, 481 × 4 displacement fields, and 936 territory × year rollups. All 10 semantic spot-checks passed. All invariants held (displacement ≤ replaceability; monotonic year ordering; band ordering; value ranges). Verdict: CONDITIONAL PASS, conditioned on one trajectory-seam issue (below) and several documentation-hygiene notes that were folded into this methodology.

Trajectory seam harmonized. The Phase 12 backward trajectory originally allowed 2025 values to sit within ±0.02 of the Phase 7 2026 anchors; Codex flagged this as a non-seamless join. The backward trajectory's 2025 values were tightened to equal the 2026 forward anchors exactly. Pre-harmonization values are preserved at phase12/reconciler/historical_trajectory_1970_2025.PRE_HARMONIZATION.json for audit. Monotonicity 1970 → 2025 confirmed post-harmonization.

Schema additions. Per Codex's recommendation that displacement_2035_moderate be directly auditable from the dataset, three new replaceability fields shipped for every occupation: replaceability_2035_moderate / _low / _high. The Phase 6 displacement_2035 value now has a matching replaceability anchor in the record. Corresponding replaceability_2040_low / _high bands were also added for band symmetry across years.

Dispersion stats on anchor-year territory records. Every territory_replaceability record at 2026, 2030, 2035, and 2040 now carries a dispersion block with standard deviation, min, max, occupation count, and the three most-exposed and three least-exposed occupations within that territory. This gives readers a shape for the distribution behind the single aggregate number.

April 21, 2026 — replaceability formula coherence fix

A coherence audit on April 21 surfaced that the stored per-occupation replaceability_2026 values were outputs of the Phase 5 blind calibrator's holistic per-occupation judgment, not the task-math formula described in this methodology. The stored values agreed with neither a strict threshold (where capability ≥ difficulty) nor any logistic crossover applied mechanically. This was a specification-implementation gap: the methodology text said task-math; the implementation produced calibrator holistic scores; Phase 10 validation didn't check formula equivalence.

The fix applied all 480 occupations under the documented logistic crossover (k=8) formula against the reconciled capability trajectory for every year 1970 → 2041, and recomputed territory replaceability as the employment-weighted mean of in-territory occupations. Phase 6 displacement values were regenerated against the new r values. Phase 4 task decompositions, capability trajectories, employment weights, barrier assignments, and territory editorial principles were all preserved unchanged — the fix is strictly in the aggregation step that was inconsistent with the documented formula. metadata.replaceability_computation in the dataset records the formula, inputs, and rationale explicitly.

Territory-level shifts were mostly modest and in defensible directions: Thinking & Leading 77 → 41 (Phase 5 calibrator had over-credited C_G saturation), Money & Data 79 → 56, Making Things 72 → 58, Buying & Selling 90 → 72. Territory shifts for predominantly physical territories were small: Land & Sea 15 → 17, Building Things 11 → 15. Occupations with REGULATORY or HUMAN_PREFERENCE barriers (lawyers, university lecturers, translators, primary school teachers) now show intermediate r26 values (28–40) that match the "capability has arrived, replacement hasn't" framing — the Phase 5 calibrator had partially embedded those barriers in the replaceability score before §2.6b correction; the recompute surfaces pure-task-math capability and lets the replacement formula handle the barriers cleanly.

Late-cycle fixes (April 20, 2026)

After the April 19 release, a dedicated spot-check surfaced a set of issues that were fixed on the dev branch before the production deploy. Documented for full transparency:

Earlier versions

v3.2 shipped April 2026 as the prior public version — editorial pass on territory assignments, display names, and barrier notes. v3.2 is preserved in archive/ai_reach_v3.2.json for lineage and rollback. Older versions (v3.0, v3.1, v2.x) are in the same archive directory. Detailed historical changelog for v3.x and earlier: see the full markdown changelog.

12. Citation and license

If referencing the dataset or methodology:

Seel, L. & Trenholm-Jensen, E. (2026). Large Labor Model (v5.0) [Dataset]. largelabormodel.com

The project is open source. Code is licensed MIT; the dataset and methodology content are licensed CC BY 4.0. Reuse, adapt, and critique freely with attribution. The full methodology markdown, changelog, and editorial process record are in the repository at docs/.

Back to visualisation Dataset on GitHub Editorial process record