Harvard's Fabrizio Dell'Acqua, Karim Lakhani, and colleagues (including Wharton's Ethan Mollick and MIT's Katherine Kellogg) ran a randomized field experiment — not a survey, an actual controlled experiment — on 758 BCG consultants, roughly 7% of the firm's individual-contributor workforce.
The results inside AI's capability zone were striking: AI-assisted consultants completed 12.2% more tasks, finished 25.1% faster, and produced output rated ~40% higher in quality.
Then came the twist. For one task deliberately designed to sit *outside* the model's reliable zone — where the AI produced convincing but incorrect analysis — consultants working without AI were right 84% of the time. Consultants with AI access dropped to 60–70%. The tool actively made them worse.
The damage came from the same behavior as the gains: trusting the output. The AI sounded right. It was fluent, well-structured, confidently stated — and wrong. There was no signal to tell you which side of the frontier you were on.