Episode #490 of Lex Fridman's podcast features machine learning researcher Sebastian Raschka and Nathan Lambert, head of post-training at the Allen Institute for AI. Their starting point is the so-called “DeepSeek moment” of January 2025, when the Chinese company released the open-weight model DeepSeek R1 with performance close to the frontier at a much lower cost, sparking accelerating competition in research and products.
Raschka argued that there is no “winner takes all” scenario, because researchers frequently move between labs and ideas circulate; what makes the difference is budget and hardware. Lambert noted the considerable hype around Anthropic's Claude Opus 4.5, while Google's Gemini 3, though powerful, is no longer discussed as much. Meanwhile, China has more than just DeepSeek: models such as Z.ai's GLM models, MiniMax's models and Moonshot's Kimi K2 Thinking are gaining ground, while Chinese companies see open weights as a way to gain influence in the US without selling expensive API subscriptions.
Names mentioned in the open-model landscape included DeepSeek, Kimi, MiniMax, Z.ai, Qwen, Mistral, Gemma, gpt-oss, Nemotron 3 and OLMo. Lambert explained that Chinese open models tend to be larger mixtures of experts and often have more permissive licenses, while models such as Llama or Gemma have restrictions, for example reporting requirements above a certain number of users. Raschka added that gpt-oss stands out because it was trained with an emphasis on tool use, which can reduce hallucinations through search or code execution rather than memorization.
On architecture, Raschka emphasized that today's models remain essentially descendants of GPT-2's transformer. The main changes are mixture of experts, multi-head latent attention, grouped-query attention, sliding-window attention and, more recently, hybrid approaches inspired by state-space models, such as Qwen3-neXt's gated delta net. These variants improve memory and inference efficiency but do not fundamentally change the autoregressive paradigm; Lambert added that significant progress also comes from systems, such as FP8 and FP4 training and faster GPU communication.
On scaling laws, Lambert explained that pre-training is not dead, but has become extremely expensive; the real cost is not so much training as serving hundreds of millions of users. Training a model can cost a few million dollars, while serving it continuously can cost billions. Raschka added that the question is where to spend compute: pre-training, mid-training, post-training or inference. Reinforcement learning with verifiable rewards and inference-time scaling brought the biggest gains in 2025, allowing models to think longer before answering.
On data, Lambert described the shift from raw Common Crawl to filtered and synthetic data; labs use optical character recognition to turn billions of PDFs into text, while rewriting sources into more structured formats. Copyright is a major issue: Anthropic lost a court case and owes authors about $1.5 billion, with the court distinguishing between purchased books and pirated copies. Raschka said human editing of code or text generated by language models has value, but also creates fatigue among open-source project maintainers.
On post-training, Lambert explained reinforcement learning with verifiable rewards: the model generates answers to math or coding problems, is scored on verifiable correctness and updates its weights using algorithms such as PPO and GRPO. This teaches it to reason step by step, correct itself and use tools. Raschka noted that a few dozen steps on a base model can send accuracy on a math benchmark soaring, not because it learns new math but because it unlocks knowledge already present; however, Lambert noted problems with data contamination, especially in Qwen models. The next phase may involve value models and process reward models to score intermediate steps.
On practical use, the three discussed tools such as Claude Code, Cursor and Codeium. Raschka prefers extensions that help him without taking over the entire project; Fridman sees Claude Code as a way to program in natural language and design at a higher level. A survey of 791 professional programmers cited by Fridman shows that experienced programmers ship just as much or more AI-generated code, and about 80% find their work more enjoyable. However, the guests warned that if we do not leave room for difficulty and personal effort, we risk never becoming true experts.
On continual learning and memory, Raschka distinguished weight updates from in-context learning; the former is currently very expensive for personalized models, with LoRA adapters offering an intermediate solution. Lambert explained that context length grows mainly through compute and data, from one million tokens toward several million. For agents, compressing the history can become an action the model learns, while DeepSeek-V3.2 uses sparse attention with a lightweight indexer to select which tokens it needs. Raschka considers tool use key to reducing hallucinations, although it does not eliminate them.
Although the discussion focused on language models, world models and robotics also came up. Raschka described world models as simulations that could also model intermediate variables, as in Meta's “Coder World Models,” rather than checking only the final answer. Lambert said he was optimistic about autonomous driving and industrial automation, but very cautious about robots in the home; Fridman emphasized that in embodied systems, safety allows almost no failures, unlike errors from a chatbot.
On AGI timelines, definitions remain unclear. Lambert suggested a practical criterion of a remote worker who can perform most digital economic tasks, and emphasized that intelligence remains “jagged”: excellent at some kinds of code, weak at others, such as distributed machine learning. The AI 2027 report, which initially predicted a superhuman programmer by 2027 or 2028, shifted toward 2031; Fridman believes it may take longer. Lambert estimated that software automation will advance this year, but automating AI research is probably more than ten years away.
On the business side, the discussion covered chatbot advertising, acquisitions and initial public offerings. Lambert said there is pressure toward consolidation, with multibillion-dollar acquisitions, but companies such as Anthropic are unlikely to be sold; meanwhile, China's MiniMax and Z.ai filed for IPOs. On Meta and Llama, the two researchers saw the effort to chase benchmarks as a source of negative publicity and internal friction. Lambert presented the ATOM initiative for US open models, arguing that research needs open weights and that the country should not depend on Chinese models; the Allen Institute for AI received a $100 million grant from the NSF for this purpose.
The discussion closed with human concerns: exhaustion in 9/9/6-style work cultures, the Silicon Valley bubble, the loss of “voice” through preference optimization and users' mental health. Lambert noted that as low-quality content proliferates, the value of physical, personal experience and trust will increase. Raschka added that AI does not remove human agency: it remains a tool that does what we ask. Despite the uncertainty, both expressed cautious optimism that people will find a way, provided we have the difficult conversations about how we want to use it.





Comments