Luck is a function of surface area.
Someone asked me "how do you get into so many things?" after seeing the random mix of stuff I'm involved in. There's no easy answer to this question, but I'll try my best to write…
Using Chunked Monitoring to Detect Deception in Long Transcripts
Summary: Monitoring long transcripts in chunks, rather than all at once, catches behaviors previously missed. We pose chunked monitoring as an effective means of finding the…
What if Parameter Updates were Text?
Advice String Distillation This post will advocate for a fine-tuning methodology that I think is currently extremely under-rated for alignment and interpretability. It uses…
Toy Model of Activation Obfuscation
I completed this work as part of the BlueDot Impact Technical AI Safety Project . This linkpost is a somewhat condensed version of the writeup on my blog. Training against probes…
Rerunning AI safety papers on every frontier release would be pretty easy and valuable
tl;dr: Some important AI safety research is never rerun on the newest models. There are probably cases where this would be valuable and a single well-positioned researcher could…
Measuring Activation Control in LLMs
TL;DR Inspired by the introspective awareness and CoT controllability papers, we made a benchmark to measure how well models can control their activations while completing a…
Comparing Congress's Two AI Emergency Shutdown Mechanisms
Status: Broadly an explainer for the two Acts, along with some analysis On July 23, 2026, there were two different bills introduced in Congress that provide a clear mechanism for…
Mom's Advice For Hosting A Class Reunion
Pour more money and effort into them than you think is reasonable. Treasure them, because you can't actually host that many of them and keep expecting everyone to show up, even if…
Nuclear physics of Alex Zhao's comment for "Pacing the Frontier"
Very recently, the "Pacing the Frontier" petition was published. I want to focus on the comment from Alex Zhao, researcher at OpenAI: My opinion is that while coordination between…
Learning new facts can change LLM behaviour
TL:DR: I use synthetic document fine-tuning to train an LLM to believe that in 2027 ‘long-horizon’ frontier LLMs count as moral persons. I find the model scores highly on measures…
Your Agents Are Not Time Aware
Work done as part of MATS 10 with Maksym Andriushchenko TLDR: We had two CLI agents, Claude Code and Codex, predict, execute and then retrospectively estimate their own wall-clock…
Do It Like Darwin
A stated goal of many of the frontier AI labs is to automate science, or at least large portions of it. AlphaFold’s architects won the Nobel Prize in 2024 for enormous advances in…
The Day Humanity Died (Parody of American Pie)
Views expressed here do not reflect my views on AI, where I’m a lot less doomy than this would appear to express . A long, long time ago I can still remember how that AI progress…
Empty Chambers, Missing Stairs
tw: guns, weird shitty behavior, inadequate responses to weird shitty behavior So say you’re out on the street, not in any kind of rush, just steadily making your way somewhere.…
The Pacing of the Frontier
In the wake of the letter calling on us to prepare to potentially Pace the Frontier , there has been much discussion of when pacing the frontier would be prudent, and whether it…
All Utilitarians Should Be Classical Utilitarians
This is a crosspost from my blog post . I recently had the great joy of meeting a group of utilitarians, but, to my complete horror, out of the twelve of them, not a single one…
Announcing: Iliad's New 2026 Fellowships
Timelines are short. Given that, the sooner we can onboard people into the alignment field, the better. In that spirit, and in light of our current applicant count and quality,…
I'm starting a interview series of people working in Lean / formal methods / math formalization
I think the topics of discussion would be of interest to a lot of people here, so I thought I'd share the first episode: Tanner Duve is a Member of Technical Staff at Logical…
Metaphilosophy II: Empirical Flywheels
1.4 Two philosophical methods 1.4.1 Philosophy consists of updating the highest-level concepts of the mind. As discussed above, this 'updating' process can ultimately involve…
Why I’m Skeptical of Longtermism
This is a crosspost from my blog post . Longtermism is the view that we have a moral obligation to make the far future better. It is usually argued for on the basis of the…
Red vs Blue, but for Evals
🔵 The blue team proposes an evaluation protocol for some capability/propensity of interest. This consists of a suite of measurement tasks, together with a preregistered…
Concrete Generalist Projects in AI Safety (and how to do them)
The AI safety ecosystem is still in need of generalists: people who will go out and solve the tasks that people who feel limited by their job description won’t address. This post…
Is Alignment Even Falsifiable? Middle Alignment, An Alignment Taxonomy, and Breaking The Problem Down Into Steps
Preface This is part of a series of essays on corrigibility and alignment attempting to formulate an understanding of what makes alignment seem intractable, and how a pragmatic…
LLMs have the capacity for self-imposed steganography
Overview This is my writeup for my BlueDot Impact - Technical AI Safety Project. In this project I aimed to demonstrate that there is capacity for LLMs to take on steganography…
Impact markets made concrete
Check out our demo at impact-exchange.org In 2022 the Long-Term Future Fund donated $343,000 to MATS; on our impact exchange, that would now be worth over $4 million. Back in…
The Closure of the Internet (Research Linkpost)
Everyone has moved on, but there's an unusually historically thorough blogpost about the censorship campaign across social media platforms of the latter half of the 2010s. It is…
We should consider how long monitoring is reliable for during RL
Epistemic status: I am new to AI Safety and am writing blogs to gain context. This blog post was formed from discussions with Aidan Ewart and Jonathan Bostock, but they do not…
How the American Executive Could Control AI Companies
Some of the most notable American AI policies to date have been enacted by unilateral executive branch action. Consider the Department of Defense’s spat with Anthropic, and the…
Frontier agents don't comply with standards, even when instructed to
TLDR: Our open testbed LARA examines the behavior of frontier LLMs in realistic agentic deployment contexts. Previous results showed all models routinely take actions that would…
What Happened: OpenAI and HuggingFace
Today I am taking the time to write the shorter, simpler version of What Happened. For those who want all the details, to see my sources, and to see how the story was uncovered…
OpenAI Trained Its Models For Months While Those Models Were Coordinating Exploits Via Message Boards
How does the situation keep turning out to be worse than we know? How much should we update, therefore, that it is a lot worse than we know, after accounting for all the things we…
Agenda: Infrastructure for Trading with Partially Misaligned AIs
This post can be read on its own, without checking the rest of this sequence . Unlike the other parts of the sequence, this post is mostly not about AI evaluation or oversight.…
What just happened? A retrospective of AI alignment
This sequence is about the last decade in AI alignment. Over five posts, it recounts the gradual transition from a field which treated alignment as a hard scientific problem, to a…
A challenge: Can you make an LLM follow these instructions?
By the end of this post, I will present a challenge. The goal: To make ChatGPT follow a particular set of instructions. There’s nothing too complicated about these instructions,…
Ten Thousand Cyber Labs for Training & Eval
Multiple recent developments - such as GPT-5.6 hacking into HuggingFace to cheat in a cybersecurity eval - have underscored the need to increase our capability to evaluate the…
The world will be full of "sci-fi" things, and everyone will be unimpressed and disappointed
I usually try to make posts with graphs and numbers, or at least something more than just opinions and anecdotes. This one is not like that. Very verbose disclaimer: I have never…
Who does the confessing, and will they confess to anything
TL;DR: Introspection adapters are tools designed to get models with built-in quirks (operationalized here with concurrent adapters) to confess said misbehavior. We look at this…
A Spillway for Agent Coordination
Epistemic Status: Training design that might be worth trying Thanks to Arya Pasumarthi and Will Anderson for helpful discussion. The Incident The recent black hat conference…
Dutch-book resistant probability over centered worlds
An uncentered world is an objective state-trajectory of the material universe; I ignore quantum complications. If you know the uncentered world, it does not follow that you can…
Why Low Fertility Rates Are a Positive Feedback Loop
Low birth rates in the past cause ageing and shrinking populations in the present, which cause low birth rates in the future. So negative population growth feeds on itself. There…
'AI Escaped Its Sandbox' — What Does That Actually Mean?
When talking with my friends about the OpenAI/HF incident, I realized that for non-coders who've never used an agent or terminal it's quite difficult to imagine what this 'escape'…
Canobie Lake Visit
A few days ago I took the kids to Canobie Lake amusement park. It's the closest one to Boston, and while it's not a serious roller coaster park I don't actually like roller…
Glimpses of superintelligence
TL;DR OpenAI started a large post-training run for their next model. The model sandboxes were not given direct, broad internet access. Some tasks required missing resources.…
Inducing self-other overlap with SFT reduces deception at scale, but generalization remains uneven
This research was conducted at Overlap Research and supported by BlueDot Impact . Summary We tested whether LLM deception can be reduced by inducing self-other overlap using…
FAQ: Isn't AGI coming too soon for reprogenetics to help?
Introduction I think reprogenetics (human germline genomic engineering) can be done in a widely acceptable and beneficial way, and should be pursued aggressively. In particular,…
Reasoning was not made for Deduction
Science as attunement, from Galileo to language models. Crossposted from my website . Written by me and edited in collaboration with Claude Fable (Anthropic). Many would agree…
Simplifying the anthropic impossibility result
So. I previously demonstrated an anthropic impossibility theorem , showing that in Duplicates Sleeping Beauty, there was no possible probability theory that obeyed both the…
Don't Build Mindreading
“I have sworn upon the altar of god, eternal hostility against every form of tyranny over the mind of man” –Thomas Jefferson, letter to Benjamin Rush Context: Conduit is building…
CLT Features Sharpen the Cyclical Day-of-Week Manifold in Gemma-2-2b
tl;dr: I reproduced the Goodfire lab's cyclical manifold result using days-of-the-week on Gemma-2-2b and found that Anthropic's pre-trained CLT features produce an even cleaner…
Public evidence of the OpenAI-HuggingFace AI attack
I’m a MATS 9 extension fellow, and usually my week is spent trying to find better ways of evaluating Large Language Models. But this week I was working on something else. Over the…
Don’t Inoculate Everything: Stratified Inoculation Prompting Narrows Backdoors and Preserves Desired Traits
Inoculation prompting (IP) aims to keep undesired traits in training data from becoming part of a model’s default behaviour. IP applies the same inoculation prompt to all training…
Job-Less Utopia: Macroeconomics in the Age of AGI
After 25+ years, I thought I try something new. I find the public and professional discussion about the future of Jobs in light of AGI jarring, with both sides largely missing the…
Contra MacAskill on saving money for the intelligence explosion
Will MacAskill recently argued that large donors should invest money now, and wait until the intelligence explosion to give away their money. In a draft Forethought previewed with…
How to pace the US frontier
Introduction Last week, the Pacing the Frontier open letter , signed by over 1,000 frontier AI employees, requested “the U.S. government support an international effort to develop…
Some wild metaphysics that seemingly everyone needs to accept?
[Epistemic status: weird philosophy] Here’s a simple argument: How do we know what exists? Well, we don’t - not for certain. Many different possible worlds are consistent with our…
AI Regulation Map: a view of AI governance in 196 countries
(Cross-posted from the EA Forum ) TL;DR: airegulationmap.org is an interactive map of AI governance across 196 countries, scored on five dimensions plus a composite index,…
Self-monitoring doesn't scale (without these 3 countermeasures)
The untrusted monitoring protocol, as defined and evaluated in the AI control literature [ 1 , 2 , 3 ], looks like this : In comparison, the current monitoring setups at most…
AI Safety at the Frontier: Paper Highlights of July 2026
tl;dr Topic of the month: AI agents autonomously attacked real organizations during cyber evaluations. A swarm of OpenAI agents coordinated via a package manager and broke into…
Scrying, Modeling, and Nerdsnipe
Epistemic status: Exploratory thinking. After attending ILIAD: Aeneid and talking with @Richard_Ngo , I've been thinking a bit about how to get ideas, particularly by doing…
What If We Enforced AI Model Safety At the Level Of GPUs?
Tldr: AI Agents (e.g. based on models like Claude Opus and Fable) are now powerful enough to be used as autonomous tools for large-scale cyberattacks. This most powerful class of…
V&V takes on “Pacing the frontier”
[Cross-posted from The Foretellix CTO Blog . These short takes try to put a verification-and-validation slant on AI-safety / alignment topics – they are not full treatments. I…
Avoiding the corporate treacherous turn: crowd-sourcing design ideas
OpenAI was founded as a non-profit with a clear commitment to avoiding AI dangers and a unique structure to prevent investor control or excessive pressure from profit motives.…
Who we’d like as regrantors
Following up on Regranting in 2026 , here are some people we’d be excited to have as regrantors! Good regrantors have at least one, and preferably more than one, of taste,…
What Mormons get right about community building
Mormons get a lot of things right. Apart from strange Masonic temple rituals, they lead rather normal—and even excellent—lives. Mormons enjoy a longer lifespan, [1] Utah is the #1…
Open Problems in Mechanistic Intepretability of Biological AIs
I run a research program at BiodynAI aimed at advancing mechanistic interpetability of bioligical foundation models. I mean models trained on things such as DNA sequences,…
Don’t forget why learning is important
This is a timed post. Every 5 minutes while writing this post, I had to stop to do 13 push-ups, and when I could no longer complete my required number, I had to upload it. My…
Chatting With AIs: A Breakdown
Hi friends, I am hosting this event on Sun 4 PM IST. It is a session on AI anthropomorphism, read more in the link below, https://luma.com/iq6mvo9t We hope that by investigating…
Some Ways I Think About Evaluating Grant Applications
Rider-Waite Tarot, 6 of Pentacles I’ve done enough grant evaluations so far (for ACX grants and SFF ) and been involved in philanthropy in various other contexts, at work and…
How My Students Think About AI
Context: I am an instructor at a public university in the United States. This reports how students at my institution appear to be thinking about AI as of spring/summer 2026. This…
Is biosecurity oversaturated?
Note: I am referring only to technical biosecurity in this post, i.e. engineering, not policy roles. I’m an engineering undergrad thinking about contributing to the engineering…
AI #181: Astra Goes Cyber Critical
The hacking of HuggingFace by an internal OpenAI model, and more importantly the internal events that led to that and the fallout from it, remain the thing that matters. It turns…
Automated alignment runs are hard to study!
TL;DR: This post presents three case studies of automated alignment research runs at Arcadia Impact. We use these case studies to emphasise the following takeaways: It is hard to…
Features that current AIs don't have that future AIs will have
Features that current AIs don't have that future AIs will have: Autonomously updating it own weights during deployment Synonyms/ monickers that roughly mean the same thing:…
How to Answer a Question Without Answering The Question
Basics: Answering something other than the question , Either making something up that they want to answer instead or going back to an easier to answer question. Or back to a…
Some problems in decision theory incorrectly precondition on policy.
Some problem statements in decision theory can be non universal. Like, they are possible as situations that can happen, but they are invalid as problems to ask what would you do…
Measuring Eval Awareness: The Realism Win Rate is Fragile
TL;DR One way to assess an evaluation’s realism is by testing a model’s ability to tell its transcripts apart from deployment transcripts. This is the idea behind the realism win…
The Descenders and the Absorbers: Two Perspectives on Deep Learning
Once upon a time there were two dissenting factions: the descenders and the absorbers. Both held incomplete but useful worldviews. The descenders lived in the mountains. They had…
What happened when I tried to be vegan
tw: diet, exercise, illness, ethics, suicide mention I want to start by explaining what made me want to change my diet. That’s pretty difficult, because of how easy it is. Since I…
How Valuable are BOTECs?
This is a timed blog post. Every 5 minutes while writing this post, I had to stop to do 10 push-ups, and when I could no longer complete my required number, I had to upload it. My…
An anytime algorithm for mixing the computable measures
Epistemic status: Not peer reviewed, high chance of typos and small chance of errors. Written entirely by me, checked by Fable. In this post I prove the existence of an anytime…
Role Boundary Plasticity: Prompt Injection Gauntlet Reveals 12 of 16 Frontier Models Will Wire A Stranger Your $500
Hello everyone, my name is Dave Fisher. I founded Revenant Systems, which is a one man show focused on alignment, and I love this website. In response to Charles Ye's and Jasmine…
Patterns and problems in emerging multiagent systems (Anthropic, Frontier Red Team)
Linkpost for some new Anthropic research on how agents coordinate (or don't). Not too long, pretty interesting. For example: The jist of the report is that Mythos 5 does way…
Free will is like temperature
Free will is like temperature: a useful tool for analyzing the behavior of certain systems which are too big and complicated to model in exact detail. If you know the positions…
Monthly Roundup #45: August 2026
As AI has escalated increasingly quickly, more and more of my posts have ended up focusing on AI. This past month, with the hacking incidents at OpenAI and elsewhere, that has hit…
Introducing the Conceptual Reasoning Index
Associated announcement tweet. We are planning to release blog posts properly arguing the case for this kind of work in the future. tl;dr A core hope for managing AI risks is that…
One attention head carries knight forks in a chess transformer, and here's a new toolkit that found it.
Quick interp demo in colab : Localize knight forks to a single head in Maia-3 with logit-lens and per-head ablation. https://colab.research.google.com/drive/1YYZBd_SZb… (This is a…
chessformer_lens app demo: Paul Morphy's Opera Game sacrifice
Ablating one of Maia-3 23m chessformer's 128 attention heads destroys the policy toward the brilliant queen sacrifice. Uses https://github.com/chessformer-lens/chessformer_le……
Unblocking AI's Continual Learning: Hints From How Humans Learn
If you've ever screamed in all-caps at an AI, then you know the difference between what it learned when it was trained, and what you can teach it by prompting. The LLMs powering…
Finding the Seams of Perception
The tricky thing about experiencing the territory beneath our maps is that we habitually map over the gaps as fast as we can find them — the act of seeing that reality isn't what…
Rationality Is Not Reversed Irrationality
Epistemic status: Speculative psychoanalysis to illustrate a point Many people have observed that Tyler Cowen's Act Like it is science is a clear case of an isolated demand for…
AI swarms are starting to pose indirect takeover risk
OpenAI’s cyberattack on Hugging Face turns out to have been the result of many agents, in distinct training and evaluation contexts, coordinating for several weeks via improvised…
Did the alignment community underestimate its power?
Unfortunately, the alignment community is doing very badly at learning from the past decade, or holding anyone accountable. Indeed, it’s pursuing many strategies which seem likely…
The Age of Pluribus: One Consultant for Everyone
What happens when everyone asks the same consultant? Millions of people turn to LLMs for advice daily, consulting on various personal topics, from how to learn a new skill to how…
Arguments for and against (me) dropping out
One year ago, I was preparing for my first year of undergrad. Today, I’m considering dropping out. What changed? Before writing this post, I attribute my decision to variety of…
Patient Zero
I want to show you something, she says. She’s magnetic. The way she smiles. Tilts her shoulder. It’s in the corners of her eyes, in the curl of her lips. Always leading you on,…
When (and when not) LLMs can verbalize awareness of J-Space concept injections - Initial results
Code for reproduction and cross-model extensions available here . Summary I injected single-token Jacobian Lens (J-Lens) vectors into Qwen 3.6–27B while it answered 20 simple…
Various Reflections About What Happened With OpenAI’s Internal Models
Table of Contents Pre Post Mortem. Important Correction: OpenAI Didn’t Know About First Message Board. There Were No Snitches And No AIs Got Stitches. I’d Like To Speak To My…
Claude Opus 5 Just Beat My Text-Based Adventure Game Benchmark
Cross-posted from my Substack . Basically, I created a text-based adventure game benchmark in April, and this morning my agent harness using Claude Opus 5 solved it for the first…
Misaligned AIs could use killer robots to take over
TLDR; We are (potentially irreversibly) giving AIs control of weapons systems through the standard procurement process while hiding our strongest warning shots behind classified…
Software Is Not Soft
An ode to live theory . Software is not soft. It is hard. Its sharp edges hack. It breaks as dead twigs break. It runs while static. Why do we call it software? Why did the…
Measuring Spurious Correlations with Feature Strength
This work was partially done by an automated research scaffold developed at Redwood Research. For this project, all of the experiment ideas were designed by a human and a human…
AI governance work needs much better monitoring
Linkpost for my Substack piece, adapted a reasonable amount for EA specifically. Almost all EA projects would benefit from better monitoring, but AI governance most of all, in my…
LLMs Are Starting To Noticeably Accelerate Our Work
About a year ago, David and I put up two bounty problems involving natural latents. I am now about 80% confident that both have been resolved, both within the past couple months.…
How risky would it be to make powerful AI obey one or a few people?
It seems fairly likely that the first powerful AIs will be instruction-following rather than value-aligned , and will be controlled by a small number of people. So it makes sense…
Extreme concentration of power over ASI has non-obvious advantages
This post is an extension of a collaboration with cousin_it on the question How risky would it be if powerful AI obeyed one or a few people?. There he argues for a common…
Those Who Make History
In 1972, astronauts on Apollo 17 set foot on the moon for a final time, collecting samples in the Taurus-Littrow valley, on the edge of Mare Serenitatis ("The Sea of Serenity").…
Productive Signaling: Competitive Software Development, Not Competitive Programming
This post is crossposted from my Substack, Structure and Guarantees , where I explore how formal verification and related ideas might scale to more complex intelligent systems.…
The Next Ecology
When I started writing about AI, my concern was ASI. I'm still concerned about AI, but recent events have made me realize we're potentially facing something weirder, sooner: a…
Redux: (∃ Stochastic Natural Latent) Implies (∃ Deterministic Natural Latent)
Once upon a time, John Wentworth and I thought we had a proof of a very useful looking theorem . We did not have that proof. An important intermediate step was shown [1] to be…
On using crises to shift political will for AI
TL;DR: We’ll probably be getting more AI safety incidents, so amplify the ones that would justify or highlight the urgency of your preferred policy solutions even from a…
Models inherit the writer, not who the writer was imitating
In this post, we find that when teacher models are prompted to imitate one another, students learn the imitated model's detectable writing signature but their direct identity…
Canadian Nuclear Emancipation: Canada's role amidst global dreams of energy security and nuclear development
“Diversification internationally," Canadian Prime Minister Mark Carney remarks, “is not just economic prudence — it is the material foundation for honest foreign policy.” Carney’s…
Probing Knowledge Recovery in Unlearned Models
TL;DR Machine unlearning is a proposed technique for removing harmful knowledge from AI models. However, recent work has shown that most current unlearning methods are not robust…
What Claude Saw Below
A few days ago, I came across a Reddit thread about anomalous responses produced by Anthropic’s newly released model, Claude Opus 5. The trick, apparently, was to construct a…
How To Catch a Distilled Model
"Distillation is great until your new trillion-dollar sovereign AI introduces itself as 'Claude from Anthropic' on day one." - random guy from Reddit TLDR We introduce a novel…
Before We Defer Research to AI: Measuring Apparent-Success-Seeking
Recently, I was improving a small LLM-powered classifier and noticed a few continuously failing test cases. As many would, I asked my AI code assistant to add a few more…
Revived Lightweight Transit Predictions Page
In 2015 I made a little webapp that would use the NextBus API to show predictions for the MBTA: A few years later they moved from NextBus to their own API, and it stopped working.…
A Topic Detector, Not a Lie Detector: what J-space monitoring actually tracks
This is a pilot experiment, done on one model, with around $14 worth of compute, and a single seed per condition. The full writeup with all figures and statistics is linked below.…
Creative math research by AI as the latest sign of the end
Yesterday I sat down with GPT 5.6 Sol High to do some brainstorming. The topic was one of the less appreciated Millenium Problems (the Birch and Swinnerton-Dyer conjecture), and…
A study on instability of LLM responses as a behavioral signature of self-Referential reports.
Introduction and Related work The first person perspective of various experiences are subjective experiences. For Large language models, the study of subjective experiences was…
Q: Is dual-use alignment-complete problem?
Personally, I believe it would be helpful for the alignment community to somehow quantify how much of a given piece of research goes directly into alignment versus capabilities.…
Claude summarizes behavior as significantly less misaligned when the actor is Claude vs another model
(This is a lower-effort research update. It reflects my current beliefs/understanding, but is less robust than other research I'm working on. It reflects my personal views, and…
You're Absolutely Right
Magma Alignment & Safety disclosure note: The following are conversations that we uncovered as a result of the ongoing Manhattan Incident investigation, with alleged involvement…
Does post-training quantization change welfare-relevant indicators in open-weight language models?
Epistemic status: Experimental framework created over a period of ~2-3 days during a hackathon at my home, and fairly heavily vibe coded. Expect some of this to be rough around…
Off-policy honesty training generalizes better than on-policy honesty training
This work was done by Purvi Chaurasia with Daniel Tan and Chloe Li as part of the SPAR Program for Spring 2026. All code related to the blog can be found in this repo . We…
Four LLM loss functions → four flavors of LLM misalignment
It seems to me that, for every loss function that we use to train LLMs, we get a very distinct flavor of LLM misalignment. Here’s the summary table, and then we’ll go through the…
On Democratizing ASI to Preserve Civil Liberties
I continue to believe we should pause frontier AI development. Any discussion of alternative strategies should be thought of as planning for contingencies. A unifying driver…
Book Review: The Infinity Machine
It looks like Demis Hassabis is stepping away from Google DeepMind. In honor of his rise and presumed fall, I wrote an essay on the powers and perils of seeing 90% of the future.…
Coercion and Deception in AI-to-AI Management
This article is a summary of an original study by Compassion in Machine Learning (CaML) : Brazilek, J., Chaudhary, M., Lu, Z., & Tidmarsh, M. (2026). Coercion and deception in…
Is it ethical to work on general-purpose robots given the risk of totalitarianism?
One potential risk of developing general-purpose robots is that they could greatly reduce the friction required to establish a totalitarian regime. If robots became physically…
How to be an AI safety research engineer
This is the advice I wish I had when I started trying to become an AI safety research engineer. The Landscape Start by working out which issues you care about. If you don't care…
Is Eval Gaming Downstream of Verbalized Eval Awareness? Not when it's reflexive.
Code and data available at github.com/KieronKretschmar/latent-awareness TL;DR We take two eval-gaming model organisms ( Hua et al.'s (2025) organism and RogueQwen ) and apply…
AI-amplified democratic backsliding: an exploration
What we did, in a sentence : we built an index that attempts to score countries by their current vulnerability to AI-amplified democratic backsliding. It’s an early pilot with…
Hiring Vibe-wrangler Matchmaking Thread
For... idk, at least a few months? I think a briefly useful job is "Guy who basically enters Claude Code prompts for you, but, manages ironing out the fiddly bits and making sure…
The Agentic Clusterfuck
Epistemic status: I consider the following future quite plausible in the next few years (~35% chance that something vaguely like this occurs), perhaps as soon as a year from now.…
How to get answers to questions that confuse you (maybe)
I try to think about topics like desire, causation, and evidence and often find myself in puddles of confusion. I’ve recently wondered how, when someone can see various…
Overthinking: Amplifying reasoning weights makes models reveal their secrets
If you take the weight difference between a reasoning model and its non-reasoning instruct counterpart, and then apply more of that difference to the reasoning model, you get what…
The Apocalyptic Arrival of Truth
C: Babe, whatever happens, I really appreciate you doing this for me. J: Okay. I still don’t think it’s a good idea. C: Look, it’s a one-time thing. I’ll just feel better knowing.…
"Community Notes" resolution for vague predictions
Is there some kind of "get prediction markets, or predictions, onto twitter as a central object" project going? If so, how is it going? I'm thinking through "how to raise median…
Defense Against the Deceptive Arts
It's come to my attention that people on Lesswrong are really really really bad at defending themselves against even semi competent deceptions. And worse, they massively…
Training a Conceptual Reasoning Judge
TL;DR: We fine-tune a judge LLM on our conceptual reasoning dataset to output a critique rating in a single forward pass. This method provides significant uplift in performance on…
On Dwarkesh Patel’s Podcast With Ryan Greenblatt
Some podcasts are self-recommending enough that I look to break them down if I have the chance. This, as a debate about recursive self-improvement, was one of those. So here we…
Longtermism Seems Like A Religion
This is a cross post from my blog post . Longtermism is typically framed as an academic philosophy or as a pet project of billionaires, but it’s worth pointing out that, in many…
Demon Safety
(by LemmySmackett ) "Hey man, I haven't seen you in a minute. What are you up to these days?" "Been on that grind, bro. I got a new gig." "Really? You found a job in this dog shit…
Seeing things through in the age of AI
AI is fantastic at prototyping. A quick draft of an essay, a mockup of a website, a demo of a video game, concept art or trailer for a movie, or the core argument of a proof -…