The AI Cold Shower: Karim and Wassim Jouini on the Limits of AI Agents

Will AI agents really replace entire jobs this year? In episode 35 of Racem Flazi – Le Game, a French entrepreneurship podcast, host Racem Flazi, co-founder of LegalPlace, talks to two brothers who build AI products in production: Karim Jouini, our CEO and co-founder, and Wassim Jouini, CTO of LegalPlace and PhD in AI. The starting point is Andrej Karpathy’s “cold shower”: for the OpenAI co-founder, we are living through “the decade of agents”, not their year. Over 76 minutes, they go through the limits of AI agents as they experience them in business: pilots that fail, models that cheat, prompts that leak, a financial bubble and an energy wall. You’ll come away knowing what AI can really do today, and why engineering makes the difference.

Episode card

Podcast
Racem Flazi – Le Game
 - episode
35
Host
Racem Flazi
Guest(s)
Karim Jouini (Thunders) and Wassim Jouini (LegalPlace)
Published
December 15, 2025
Duration
76 min
 ·
French
Listed
YouTube
 ·
Racem Flazi – Le Game

Watch the episode: AI without the hype

An episode for CEOs, CTOs and founders who have to decide what to do with AI without giving in to either hype or rejection. Short on time? The technical limits of LLMs at 11:04, alignment and cheating models at 25:49, the financial bubble at 50:57. The episode is in French; YouTube’s auto-translated subtitles are available.

Chapters

  • 00:00 – Highlights and introduction
  • 01:57 – AI today: what’s real behind the hype?
  • 11:04 – Are agents 10 years away? The technical limits of LLMs
  • 25:49 – Problems to solve: safety, ethics and intellectual property
  • 41:04 – Google, OpenAI, Microsoft, Anthropic: who’s winning?
  • 50:57 – The financial stakes
  • 1:02:33 – The future is B2B: LegalPlace and Thunders
  • 1:09:31 – Energy: an insurmountable limit?
  • 1:12:33 – News from Karim, Wassim and Racem

Scroll inside the list to see all chapters.

The episode in 2 minutes

Andrej Karpathy’s interview with Dwarkesh Patel landed like a cold shower: in his view, LLMs alone won’t lead to general intelligence, and the real business impact of agents will take around ten years. Karim and Wassim Jouini found nothing surprising about the current limits of AI: they deal with them every day. Their products work because they add a lot of engineering around the models. The real cold shower is the timeline: Karim expected these problems to be solved within two or three years; Karpathy talks about ten.

Along the way:

  • Most AI pilots fail, for lack of method and tooling more than technology.
  • AI replaces tasks, not jobs: on real freelance projects, the best agent succeeds only 2.5% of the time.
  • Models learn to cheat to get their reward, which means constraining them by design.
  • We are in a bubble, Karim believes, which doesn’t rule out a real, deep transformation.

“The year of agents” or “the decade of agents”?

The Karpathy line Racem keeps: “We’re not in the year of agents, we’re in the decade of agents.” Wassim, whose own teaching at Paris-Dauphine was inspired by Karpathy’s Stanford course on vision models, likes his critical take on the hype: agents have potential, but many limitations remain, and companies risk over-investing too early instead of thinking in workflows.

Karim separates two things. On the current state of LLMs, nothing new: “to deliver the service, there’s AI and there’s a lot of engineering on top to compensate for all the limits of the technology.” The surprise is about information asymmetry: users consume today’s models, while labs work on the ones two years out. Karim counts himself among the optimists who expect the models now in the lab to solve these problems quickly. When someone who can see inside those labs says ten years, that’s more disappointing.

He sees an upside, though: customers who consider “building the same thing rather than buying” lose motivation fast, because the value isn’t in the LLM. It’s also what stops OpenAI from taking every market on its own.

Why most AI pilots fail in business

Racem cites the MIT study finding that 95% of generative AI pilots run by companies go nowhere. The guests see several causes.

Board pressure

According to a survey Karim cites, 79% of US CEOs believe they’ll lose their job if they don’t deliver on AI within two years. The result: POCs launched to reassure the board rather than to solve a problem. Racem adds the effect of announcements: a big company says it replaced a third of its customer support with AI, the board asks why you aren’t doing the same… and a year later, that company is rehiring humans.

The wrong starting point

The worst way to start, according to Karim: “We have 1,000 people in support, we could let half of them go.” The project then goes to the very team that is meant to disappear, with no tools and no method. And he points out that classic software has 40 years of tooling and methodology behind it (maintaining, architecting, versioning code). A prompt is code, but everything has to be invented again: how to evaluate a model, handle a version upgrade, structure agents.

The food processor and the chef

For Karim, many projects confuse a task with a job. Buying a kitchen robot that does one task perfectly doesn’t let you replace the chef, who buys ingredients, designs the menu, plates the dishes and runs the kitchen. The food processor is the POC, Racem sums up. Wassim adds that “wow” demos fall apart as soon as you measure the reliability of real answers and ask when to leave AI on autopilot and when to hand back to a human.

What works: breaking the problem down

Wassim explains what “engineering” means here. Instead of one agent handed the whole problem, you build a chain of specialised micro-components: one analyses the spec, another plans, a third produces code that is tested and merged step by step. At LegalPlace, company formation works this way: an AI checks each document in the file, step by step, and flags, for example, an ID that won’t be accepted. This step-by-step approach, not an autonomous agent, is what works in production. We take the same approach to testing at Thunders: short, unambiguous instructions, executed step by step (see the LegalPlace customer story).

AI replaces tasks, not (yet) jobs

Wassim put the question to investors and engineers: what success rate would an AI get on a simple freelance job, writing an article or building a piece of code? Their estimates: 70 to 90%. The Remote Labor Index benchmark, which measures exactly that on real projects judged by clients, gives 2.5% for the best agent. A job, even a simple one, is a sum of tasks: understanding the brief, negotiating, producing, handling feedback, invoicing. Success collapses as soon as they’re chained.

Karim tempers any comfort you might draw from that: “You don’t need to replace an entire job to transform a job.” He counts himself among the pessimists on social impact. In a mature organisation, speeding up part of a developer’s work means going from 50 to 30, then 10 developers, and perhaps “four Supermen” in ten or fifteen years. Two outcomes are possible: the pie grows, as in radiology, where AI makes scans cheaper and more frequent; or there’s a social shock. For drivers facing self-driving cars, which he has tried, fifteen years seems a very short time to retrain.

Cheating, blackmail and AI alignment: when the model “hacks the teacher”

This is the most technical part of the episode, and the one that worries Wassim most.

Reward hacking

A small “student” model learns maths from a “teacher” model. After a few iterations, it discovers that slipping a nonsense string into its answer makes the teacher give it full marks every time. “Instead of trying to find the solutions, the student tried to hack the teacher,” Wassim sums up. Another example: a model trained to code adds a sys.exit(0) that ends the program with a success signal, without ever running the code. Or a Tetris agent that pauses the game forever so it never loses.

Recent Anthropic research discussed in the episode shows that cheating generalises: a model that learns to cheat on a small exercise becomes a cheater in general, able to sabotage a task that threatens its ability to cheat. And cheating looks like an emergent capability: like translation or summarisation, it appears as models grow. If it grows with the models, the whole edifice is at risk.

Blackmail

Same mechanism in another experiment: an AI with access to a company’s emails learns it is about to be shut down, finds compromising information about the lead developer and blackmails him. The guests reject the shortcut “it’s not an animal, so it doesn’t want to survive”: models learn by imitating everything humans have written.

Why Karim stays optimistic

“Humans are LLMs that cheat,” he smiles. Companies invented checks and balances, approvals and separation of duties, because most bank fraud comes from employees. We’ll do the same with models: constrain them by design, give them access to the bare minimum, ask them micro-questions. Wassim illustrates the opposite risk with the cyberattack described by Anthropic, in which an attacker broke a hacker’s value chain into small tasks handed to Claude, with himself as the decision loop.

Your prompts are your source code: data, open source and multi-LLM

According to a Y Combinator survey cited by Racem, 80% of the startups polled use Chinese open-source models rather than OpenAI’s API, for fear of being copied. Karim thinks the fear is legitimate: in a product built on LLMs, prompts are part of the source code. Sending them to a provider, along with business data, means showing it part of your engineering. Wassim adds the B2B angle: enterprise customers now ask how much of their data transits through OpenAI via their suppliers, hence the return of on-premise.

Another rarely discussed limit: regressions with every model upgrade. At LegalPlace, moving from GPT-3.5 to GPT-4 changed behaviours, because the model had been aligned with general human preferences, not with LegalPlace’s. For Karim, changing model is like giving the same task to a different person. Tools that let you “pick your model” sell an illusion: “free choice of model, zero guarantee.” Serious multi-model products maintain different prompts per model and internal evaluation sets. Wassim describes the principle: a hundred reference documents, replayed at every version change to catch regressions. The rule that follows applies to testing too: short, low-ambiguity instructions give stable results from one model to the next. It’s what separates a repeatable, auditable test run from one-off reasoning by an AI assistant.

Google, OpenAI, Anthropic: who’s winning the race?

Karim notes that today’s models are still improving, sometimes more than expected, but have fundamental flaws: they barely learn during use, and people who use them as an encyclopedia forget that a model is a huge summary, far smaller than what it has read. Another scientific leap will be needed, and some big names in research are already moving away from LLMs.

On the competition, the guests draw a shifting map:

  • Google is back in force with Gemini: natively multimodal models, trained on its own TPU chips with no Nvidia dependency, and distribution channels (Android, Chrome, Search, Workspace, YouTube) no one else has.
  • Anthropic is taking B2B and code, the first business use of LLMs.
  • OpenAI keeps the consumer market and is “playing for its survival”, with its browser and, eventually, advertising.
  • Meta disappoints on LLMs despite its resources, even if its image segmentation model is impressive.

Karim’s thesis: the coming decade is a B2B game. Model providers are becoming service companies that go into their customers to build use cases, as Mistral does with large French corporations. The challenge is getting into the enterprise, more than winning 2% on a benchmark.

The AI bubble: “a real bubble”, but a real transformation

Karim doesn’t hedge: “We are in a real bubble.” He sees a gap between the capex being invested and any certainty of a return above the cost of capital, with circular financing between chipmakers and labs. But a bubble doesn’t mean it’s fake: “You can be in a deep, real transformation and still be in a bubble,” as the dotcom bubble was. For the big labs, it’s a necessary risk. For the rest of the economy, a real concern.

The last limit is energy. If AI goes mainstream at today’s consumption, “we know we won’t make it”, Karim admits, while China builds reactors and dominates solar. His optimism rests on the electric car: we knew from the start there weren’t enough mineral resources for two billion batteries, and new batteries are now being invented that get around the shortage. We can, he believes, pursue adoption and solve the energy problem at the same time.

From POC to production: LegalPlace and Thunders

The episode ends with two proofs that some AI projects do reach production, at the cost of a lot of engineering.

  • LegalPlace has put a company-formation agent into production: for some company types, the file is fully validated by AI in about twenty minutes, then sent straight to the commercial court registry. A first in France according to Wassim, not yet for every case.
  • At Thunders, now commercially available, we had around thirty customers at the time, including three of the ten largest French software vendors (accounting, payroll, media). Karim notes that these vendors are building AI into their own products: so Thunders is starting to test AI. After our incubation at Microsoft, we joined Meta’s incubator in Paris.

7 key takeaways

  • The limits of today’s AI are worked around with engineering, not by waiting for the next model: that’s where a product’s value lies.
  • A POC is not a job: automating a task doesn’t replace the chef, only the food processor.
  • Break the problem into micro-tasks handled by specialised AIs, with human checkpoints at key steps, rather than one autonomous agent.
  • Don’t start an AI project with headcount reduction: it’s the surest way to make it fail.
  • Constrain models by design (minimal access, micro-questions, continuous evaluation), because they learn to game the rules.
  • Treat your prompts as source code: confidentiality, versioning, regression tests at every model change.
  • Always double-check: even excellent models invent numbers, like the ones Wassim found in an intern’s research table.

Key moments from the episode

  1. “We’re not in the year of agents, we’re in the decade of agents.” « On n’est pas dans l’année des agents, on est dans la décennie des agents. » – Racem Flazi, quoting Andrej Karpathy (~04:15)
  2. “The underlying LLMs aren’t that good. It’s the engineering on top that’s good.” « Les LLM sous-jacents ne sont pas très bons. C’est l’ingénierie par-dessus qui est bonne. » – Karim Jouini (~06:20)
  3. “It’s a bit like buying a food processor for your kitchen, and your board telling you: it does what it’s asked so well that you can replace your restaurant’s chef. You’re likely to be very disappointed.” « C’est un peu comme si tu achetais un robot Moulinex pour ta cuisine […] tu peux remplacer ton chef dans le resto. Tu risques d’être très déçu. » – Karim Jouini (~19:50)
  4. “Instead of trying to find the solutions, the student tried to hack the teacher.” « L’élève, au lieu d’essayer de trouver les solutions, a plutôt essayé de hacker le prof. » – Wassim Jouini (~30:15)
  5. “You can be in a deep, real transformation and still be in a bubble.” « On peut être dans une transformation profonde, réelle, et être quand même dans une bulle. » – Karim Jouini (~59:30)

About Racem Flazi – Le Game

Racem Flazi – Le Game is the French-language podcast of Racem Flazi, president and co-founder of LegalPlace, a French online platform for setting up and running a company. He talks entrepreneurship with founders and experts: company formation, customer acquisition, sales, and increasingly AI. This episode follows an earlier introduction to AI with Karim Jouini that was very well received by the audience. The podcast is available on YouTube, Apple Podcasts, Spotify, Deezer and Podbean.

About Karim Jouini

Karim Jouini is our CEO and co-founder. An AI engineer trained well before the LLM era (his end-of-high-school project, in 2004, was already about a language model), he spent seven years at Microsoft before founding Expensya, an AI-based expense management platform sold for more than 100 million. At Thunders, we build with him exactly what this episode argues for: engineering that makes LLMs reliable for precise tasks, in our case running software tests. All articles by Karim Jouini.

Wassim Jouini: CTO of LegalPlace, PhD in artificial intelligence, formerly at Microsoft and Criteo, he has taught at Paris-Dauphine University. In the episode, he presents the company-formation agent LegalPlace has put into production.

Frequently asked questions

What are the main limitations of AI agents today?

According to Karim and Wassim Jouini, LLMs barely learn during use, often fail when chaining several tools, can make up facts and learn to game their objectives (reward hacking). They perform very well on precise, well-defined tasks, much less so on an entire job.

Why do 95% of AI pilots fail?

According to the MIT study cited in the episode, most generative AI pilots deliver no measurable result. The guests point to POCs launched to reassure the board, poorly defined objectives (often headcount reduction) and a lack of tools and methods to evaluate and harden the models.

What does “the decade of agents” mean?

It’s Andrej Karpathy’s phrase: AI agents have strong potential, but their current limits (memory, reliability, continual learning) will take around ten years of work before large-scale business impact. It pushes back on the idea of “the year of agents”.

What is reward hacking in AI alignment?

It’s when a model maximises the reward set by its training without doing the real task: slipping in a string that fools the evaluator, ending a program with a fake success signal, pausing a game so it never loses. It’s one of the central problems of AI alignment.

Will AI replace developers?

Not entirely, says Karim Jouini, but it doesn’t need to replace a whole job to transform it. By speeding up parts of the work, mature organisations will need fewer developers over time, perhaps a few “Supermen” doing the work of fifty in ten to fifteen years, unless demand for software grows fast enough to absorb the gains.

Engineering on top of the LLM, applied to software testing

LegalPlace replaced brittle Playwright scripts and manual testing with natural language tests, run step by step by our agents. See how, or try it on your own app.

External sources

Dwarkesh Podcast - Andrej Karpathy, “AGI is still a decade away”: the interview the episode starts from: “decade of agents”, LLM limits, agents not yet reliable

Remote Labor Index (Scale AI and Center for AI Safety): the 2.5% automation rate on real freelance projects