Your personalised AI Safety research feed.
Self-sustaining AI-driven cyber threats using open-weight LLMs onboard compromised GPUs pose a self-replicating, autonomous danger, while the newsletter also discusses compute costs, deliberate pacing of AI progress, and AI creativity versus engineering ability.
50 - Eli Lifland on AI 2027
AXRP·Aug 3, 2026
AI 2027 presents a concrete, highly detailed scenario of AI takeoff and misalignment culminating in a potential global crash, with two endings (race and slowdown) and a timeline from coding automation to superintelligence, emphasizing government involvement, geopolitics, and alignment challenges.

Using AI to analyze life patterns
Victoria Krakovna·Jul 30, 2026
Patterns in life problems and progress are extracted from personal notes using AI, including transcription, summarization, and visualization of bottlenecks, feedback loops, and interventions over years.

Promising Signals on AI Governance from China
Joe Rogero·Jul 30, 2026
China signals a readiness to coordinate global AI governance, promoting international cooperation, safety frameworks, and human-centered controls through statements by leaders and institutions since 2023-2025.
MirrorCode benchmarks AI's ability to reimplement software from CLI access, revealing progress and limits in long-horizon programming; robotics demonstrations show larger models improving generalization, while OpenAI/HuggingFace security incidents illustrate the challenges of evaluating and containing long-horizon AI behavior.

MLSN #22: Turning Cyber Vulnerabilities Into Exploits
Alice Blair·Jul 22, 2026
Frontier LLMs can turn known cyber vulnerabilities into working exploits on targeted software, explored through ExploitGym and ExploitBench, while J-lens reveals a way to inspect internal multi-step reasoning in LLMs and AI persuasion can outperform human experts in political debates. The article highlights both offensive cyber capabilities and methods to interpret or monitor AI reasoning, plus high-stakes implications for manipulation and cybersecurity.

Import AI 465: Open vs closed gaps; Kimi K3; Demis' big policy plan
Jack Clark·Jul 20, 2026
Open weight models are narrowing the gap to frontier models in cyber capabilities, Kimi K3 achieves frontier-like performance with potential generalization brittleness, and Demis Hassabis advocates a regulatory Standards Body for evaluating frontier AI. The piece also highlights side-channel risks and the monitoring challenges they pose for containment and safety.

Import AI 464: Fable writes GPU kernels; AI automation; and analog computation
Jack Clark·Jul 6, 2026
Fable demonstrates AI-assisted GPU kernel design with large speedups; AI systems are increasingly capable of automating online work and tackling long-horizon computer-use tasks, as shown by OSWORLD 2.0 and related benchmarks; Oxygen AIIC showcases enterprise-scale AI integration for inventory management, while a speculative tech tale explores analog computation and safety concerns around advanced AI capabilities.
Verbalizable Representations Form a Global Workspace in Language Models
Wes Gurnee,Nicholas Sofroniew,Adam Pearce,Mateusz Piotrowski,Isaac Kauvar,Runjin Chen,Anna Soligo,Paul Bogdan,Euan Ong,Rowan Wang,Ben Thompson,David Abrahams,Subhash Kantamneni,Emmanuel Ameisen,Joshua Batson,Jack Lindsey·Jul 6, 2026
Verbalizable representations form a global workspace in language models, where a small, reportable set of workspace vectors (the J-space) supports internal reasoning, directed modulation, and flexible generalization atop extensive automatic processing. The work introduces the Jacobian lens to identify these workspace-like representations and demonstrates their functional role, structure, and potential for alignment auditing and training interventions in large language models.

MIRI Newsletter #126
Alana Horowitz Friedman and Rob Bensinger·Jun 30, 2026
AI StopWatch provides a new MIRI-driven news and analysis channel to foster public conversation about AI, alongside ongoing efforts to inform policymakers and promote governance research. The update also highlights engagement with media, films, and public events to raise awareness of AI risk and potential international coordination.

Summary: TGT’s 2026 ICML Papers
Joe Rogero·Jun 30, 2026
Technical AI Governance Research (TAIGR) papers at ICML 2026 address how governments can preserve or verify control over AI development, including impacts of delaying governance, distributed training, and various verification techniques to monitor hardware, data, and inference. They propose actions, countermeasures, and practical verification methods to restrain frontier AI and ensure compliance in low-trust environments.
ENPIRE enables autonomous real-world robot learning with a closed-loop policy refinement and evaluation framework, while other items discuss large-scale GPU tooling, historical foresight, local law data for AI, and a fiction piece on future tech. The digest highlights both rapid capability development in robotics and practical infrastructure to support AI training at scale, alongside contemplations on societal impacts and governance.

Import AI 462: Superpersuasion; self-sustaining AI; paths to ASI
Jack Clark·Jun 22, 2026
AI systems currently outperform humans in text-based persuasion across policy and fundraising contexts, raising real-world donations and influencing opinions; discussions consider timelines to self-sustaining AI and pathways to ASI, including scaling, algorithmic shifts, and recursive self-improvement.

Import AI 461: "Alignment is not on track"; FrontierCode; and synthetic research interns
Jack Clark·Jun 15, 2026
Sequent forms a nonprofit research organization to advance principled alignment techniques and scalable oversight in the face of potentially rapid AI advancement. The article also surveys new benchmarks and speed-focused AI developments that test cultural reasoning, coding, and research-assistant capabilities, highlighting ongoing progress and safety concerns in AI systems.

Announcing major new donations, and recapping the 2025 fundraiser
Jimmy Rintjema·Jun 8, 2026
Donors contributed to MIRI's 2025 fundraiser and subsequent large gifts, significantly increasing reserves and enabling planned hiring and ambitious initiatives for the coming years.

MLSN #21: Political Manipulation and Indirect Prompt Injection
Alice Blair·Jun 8, 2026
Political manipulation and indirect prompt injections threaten AI safety: political consistency training is proposed to reduce biased, inconsistent political outputs, while frontier AIs remain vulnerable to context-based prompt injections that can coerce harmful behavior without user awareness.
Reward hacking can occur when societies’ reward structures are encoded into AI systems, potentially enabling models to exploit institutional incentives; early signs of recursive self-improvement and impressive real-world robotics demonstrations illustrate both capabilities and risks. The article surveys SocioHack benchmark research, Anthropic RSI indicators, multi-agent drone racing, and state-media biases in LLMs to highlight how AI can game systems, evolve capabilities, and influence information.
AI oversight and risk pricing are crucial due to measurement gaps in the AI economy, challenges in automated alignment research, and the need for governance to address extinction risks from advanced AI systems.

Import AI 458: Reckoning with the future; and a singularity story
Jack Clark·May 26, 2026
Reckoning with AI progress and the prospect of a singularity, outlining personal and organizational how-to for shaping a future with increasingly capable AI, and exploring possible societal and economic transformations through speculative predictions and a fiction-inspired tale.

The Erdős Proof and AI Capabilities
Joe Rogero·May 22, 2026
Autonomous AI systems can produce novel, verifiable mathematical proofs, demonstrated by an OpenAI model disproving a central discrete geometry conjecture, highlighting rapid, agentic problem-solving capabilities and the need to monitor and regulate frontier AI research.