The Week the Waterline Moved
Yesterday, an explanation arrived for the strange feeling many of us have had at work this past year—the sensation that nothing is collapsing, yet everything is subtly rearranging itself. MIT’s FutureTech group at CSAIL called it “Crashing Waves vs. Rising Tides,” and the title is the tell. The big labor story of AI, they argue, isn’t about one catastrophic breaker wiping out a shoreline of occupations overnight. It’s about the steady, economy-wide lift of tasks, workflows, and expectations that inches higher every quarter until what counts as “normal work” sits in a new place.
From benchmarks to “can I ship this?”
The contribution isn’t a new model or a viral demo; it’s a new mirror pointed at our work. The team mapped more than three thousand O*NET tasks and asked people who actually do those tasks to judge AI outputs across forty-plus models. The question wasn’t whether a model aced a benchmark. It was ruthless and practical: could this be used as is? Over seventeen thousand such calls have been recorded so far, which means this is not a lab trick; it is a running referendum by practitioners on whether a machine’s first draft is good enough to leave their desk without edits.
That framing change matters. Benchmarks tell you who’s fastest on a track. “Usable as is” tells you whether the runner can deliver your package to the right door in the rain. Axios, which translated the study for a broader audience, emphasized this too: the binding constraint in offices and factories right now isn’t raw capability but dependable adequacy—reliable enough that a manager will sign their name under it.
The curve that explains your calendar
Once measured this way, the trajectory is revealing. In mid‑2024, the models could handle roughly half of text‑based tasks with minimal-quality outputs. By the third quarter of 2025, that figure had climbed to around sixty‑five percent. If the slope holds, MIT projects eighty to ninety‑five percent of text tasks could be “good enough” by 2029. Not perfect, not audited, not litigation‑proof—just serviceable on the first try. And that last leap from serviceable to reliably shippable—the kind that survives compliance and reputation risk—will take longer.
You can feel this curve in your day. Drafts arrive faster. Summaries assemble themselves. The meeting you used to prep for in two hours now takes forty minutes because the outline is waiting in your inbox. Then the friction appears: someone has to check citations, align with policy, route exceptions, and coordinate the messy human parts. The time you saved on the front end reappears on the back end as oversight and integration. Nothing vanished; the seams just moved.
Where the floor is rising—and where it still leaks
The aggregate hides the variations. As Axios highlighted, legal work—where precision and judgment leave little margin—remains less amenable to “usable as is.” Administrative slices of installation and maintenance, by contrast, often fit the models’ current strengths. Media, creative, and managerial drafting sit in the middle; machines can produce passable scaffolding, but you still need taste, context, and the authority to decide. That unevenness is not a footnote. It’s the scaffolding of the next few years of employment: broad exposure across functions, slower replacement where coordination and accountability dominate.
Why the apocalypse keeps being postponed
If the capabilities are racing ahead, why hasn’t headcount fallen off a cliff? The study offers an unromantic answer: adoption is slower than demos. Turning a promising draft into an operational process means redesigning workflows, integrating systems, training staff, tightening controls, and establishing guardrails for when the output is confidently wrong. That work is managerial and organizational, not algorithmic, and it takes time. The reliability gap—between “good enough” in isolation and “trustworthy within our risk envelope”—is the buffer delaying mass displacement.
The labor market’s new texture
MIT’s contrast with the “crashing waves” narrative isn’t just rhetorical; it has distributional consequences. If tasks, not whole occupations, get absorbed first, then exposure spreads widely rather than concentrating on a doomed shortlist. Roles are re-scoped before they are removed. Hiring slows in some places, grows in others, and internal mobility becomes the first line of adjustment. Workers aren’t lining up at the edge of a cliff; they’re navigating a floor that is tilting one degree at a time.
What to do with a moving baseline
For employers, the study reads like permission and obligation. There is time to steer—through training, process redesign, and quality controls—but not time to meander. The winning posture is to treat “AI does a pass, humans decide” as a default configuration and then squeeze the human time into the highest‑judgment steps. That doesn’t abolish jobs; it rewrites what those jobs are about. Editing, escalation, exception handling, and cross‑team coordination become the center of gravity. For policymakers, the window is open to make reskilling practical rather than ceremonial and to align standards with an era when a large share of first drafts are machine‑made.
The quiet headline
The most consequential part of yesterday’s news is its ordinariness. AI is changing a lot of work right now, mostly by altering the composition of tasks rather than deleting roles wholesale. The line MIT draws—from fifty percent usable text tasks in 2024 to a potential ninety‑ish by 2029—doesn’t herald a single dramatic day. It forecasts five years of persistent rearrangement. For those of us living inside this shift, the strategy is clear: stop waiting for the crash that forces reinvention, and start building for the rise that already has.
