Skip to module content
Module 06 · ~9 min

Humans on the Jagged Frontier

The largest field experiment on AI and knowledge work — huge gains, a hidden cliff, and three working styles.

Reading progress
0/7 · 0%

The big idea

💡Key idea
AI capability isn't a smooth surface — it's a jagged coastline. Some tasks sit well inside the model's reliable zone; others sit just outside it. The dangerous property: output quality looks identical on both sides. Fluent, well-structured, confident — and sometimes completely wrong. Adoption design is therefore frontier management: map which tasks sit where, build structural (not aspirational) oversight at the boundary, and re-map every time the model upgrades. This is the human-side twin of Volume 1's eval discipline. Evals locate the frontier empirically for your system; this section is what you do organizationally with that map.
Quick check
1 question · instant feedback
0/1
  1. The core danger the 'jagged frontier' names:

Numbers that matter

+12.2%
more tasks completed by AI-assisted BCG consultants on inside-frontier work
+40%
higher rated quality, 25.1% faster — inside-frontier tasks
84% → 60–70%
accuracy drop on tasks designed to sit outside AI's capabilities when consultants used it
Quick check
1 question · instant feedback
0/1
  1. The three working modes identified in the studies:

Deep dive

3/3 open

Harvard's Fabrizio Dell'Acqua, Karim Lakhani, and colleagues (including Wharton's Ethan Mollick and MIT's Katherine Kellogg) ran a randomized field experiment — not a survey, an actual controlled experiment — on 758 BCG consultants, roughly 7% of the firm's individual-contributor workforce.

The results inside AI's capability zone were striking: AI-assisted consultants completed 12.2% more tasks, finished 25.1% faster, and produced output rated ~40% higher in quality.

Then came the twist. For one task deliberately designed to sit *outside* the model's reliable zone — where the AI produced convincing but incorrect analysis — consultants working without AI were right 84% of the time. Consultants with AI access dropped to 60–70%. The tool actively made them worse.

The damage came from the same behavior as the gains: trusting the output. The AI sounded right. It was fluent, well-structured, confidently stated — and wrong. There was no signal to tell you which side of the frontier you were on.

The research identifies two productive working modes and one unproductive one.

**Centaurs** divide the labor cleanly — they keep strategy, judgment, and interpretation for themselves, and delegate structured production work (drafting, summarizing, formatting) to the AI. The division is deliberate.

**Cyborgs** interleave with the model continuously throughout a task — thinking with it, not just assigning to it. The boundary between human and AI contribution is fluid.

The 2025 follow-up (HBS Working Paper 26-036, n=244) added a third mode: **self-automators**, who hand over the entire task — including the oversight. They treat the AI output as the final product.

Both productive modes require genuine expertise. You cannot validate output on a frontier you can't locate yourself. This yields the study's most strategically important implication: **expertise becomes more valuable under AI, not less.** Validation — knowing when to trust and when to push back — replaces production as the scarce, high-value skill.

The natural response to self-automation risk is to tell people to 'review carefully.' It doesn't work reliably.

The 2025 follow-up found that disciplined intention alone doesn't prevent consultants from effectively outsourcing their judgment to the model. When the output looks good, review becomes confirmation, not evaluation.

The fix is structural: design the workflow so that independent human analysis happens *before* AI consultation — not after, as a check on AI output. Or run parallel validation (human and AI reach conclusions independently, then compare). Or separate the role of AI-assisted drafter from the role of final decision-maker, so one person can't self-automate both.

Human-in-the-loop is an architecture, not an intention. Build it into the process, or it won't hold.

Quick check
1 question · instant feedback
0/1
  1. Correct implication for expertise:

In the field

🔬Worked example
Frontier map for one workflow, three columns: (1) confidently inside frontier — automate with sampling QA; (2) boundary zone — structural human validation (independent human analysis before AI consultation, or parallel validation, or separated drafting-vs-judgment); (3) outside frontier — human-led, AI-assisted drafting at most. Wire review gates from Volume 1 §5 to match the map; schedule re-mapping into the quarterly eval refresh.
🚫When not to reach for it
Static frontier maps rot fast — last year's 'AI can't do X' hardcoded into policy after two model generations moved the coastline is a fresh failure mode. Aspirational 'humans will review carefully' policies fail the self-automation test; oversight must be structural (workflow-designed independent analysis, parallel validation, separation between drafting and final judgment), not intention-based.
Quick check
1 question · instant feedback
0/1
  1. Preferred oversight design after the 2025 follow-up:

Pitfalls & takeaways

Failure modes

  • Uniform rollout. Giving everyone the tool with no map of which tasks sit inside vs outside — guaranteeing both under-use on safe tasks and over-trust on dangerous ones.
  • Approval-gate theater. A human 'reviews' every output but has no independent basis or time to validate — self-automation with a signature line.
  • Junior-first deployment. The people least able to locate the frontier get the most exposure to it; the skill-building pipeline that creates validators quietly erodes.
  • Static frontier maps. Last year's 'AI can't do X' hard-coded after two model generations moved the coastline.

Durable takeaways

  • The frontier is a coastline, not a wall — invisible and moving.
  • Fluent + confident + sometimes wrong is the dangerous property.
  • Expertise appreciates under AI (validation is scarce), production commoditizes.
  • Structural oversight beats disciplined intention every time.

Do the work

📦Artifact to produce
Frontier map (3-column table) for every solution-design deliverable, refreshed quarterly with evals.

Sources

  • · Dell'Acqua, Lakhani, Mollick, Kellogg et al., 'Navigating the Jagged Technological Frontier' (2023) — n=758 BCG consultants
  • · HBS Working Paper 26-036 (2025) — n=244 BCG consultants