📰 NewsWatchdog.com — Less Wrong

🐕
Aug 16, 2026 at 18:04 UTC · ← All Sources · ← Main Feed
Less Wrong145
Less Wrong 21h ago

Luck is a function of surface area.

by sid.the.manne@gmail.com

Someone asked me "how do you get into so many things?" after seeing the random mix of stuff I'm involved in. There's no easy answer to this question, but I'll try my best to write…

⚡64 · 🛡100
Less Wrong 21h ago

Using Chunked Monitoring to Detect Deception in Long Transcripts

by sfereido

Summary: Monitoring long transcripts in chunks, rather than all at once, catches behaviors previously missed. We pose chunked monitoring as an effective means of finding the…

⚡58 · 🛡100
Less Wrong 21h ago

What if Parameter Updates were Text?

by DaemonicSigil

Advice String Distillation This post will advocate for a fine-tuning methodology that I think is currently extremely under-rated for alignment and interpretability. It uses…

⚡53 · 🛡100
Less Wrong 1d ago

Toy Model of Activation Obfuscation

by Jesse Li

I completed this work as part of the BlueDot Impact Technical AI Safety Project . This linkpost is a somewhat condensed version of the writeup on my blog. Training against probes…

⚡43 · 🛡100
Less Wrong 1d ago

Rerunning AI safety papers on every frontier release would be pretty easy and valuable

by Zephaniah Roe

tl;dr: Some important AI safety research is never rerun on the newest models. There are probably cases where this would be valuable and a single well-positioned researcher could…

⚡41 · 🛡100
Less Wrong 3d ago

Measuring Activation Control in LLMs

by Marek Kowalski

TL;DR Inspired by the introspective awareness and CoT controllability papers, we made a benchmark to measure how well models can control their activations while completing a…

⚡38 · 🛡100
Less Wrong 2d ago

Comparing Congress's Two AI Emergency Shutdown Mechanisms

by Philip Dowdell

Status: Broadly an explainer for the two Acts, along with some analysis On July 23, 2026, there were two different bills introduced in Congress that provide a clear mechanism for…

⚡36 · 🛡100
Less Wrong 1d ago

Mom's Advice For Hosting A Class Reunion

by jenn

Pour more money and effort into them than you think is reasonable. Treasure them, because you can't actually host that many of them and keep expecting everyone to show up, even if…

⚡34 · 🛡100
Less Wrong 1d ago

Nuclear physics of Alex Zhao's comment for "Pacing the Frontier"

by Dante Dam

Very recently, the "Pacing the Frontier" petition was published. I want to focus on the comment from Alex Zhao, researcher at OpenAI: My opinion is that while coordination between…

⚡32 · 🛡100
Less Wrong 1d ago

Learning new facts can change LLM behaviour

by Richard Juggins

TL:DR: I use synthetic document fine-tuning to train an LLM to believe that in 2027 ‘long-horizon’ frontier LLMs count as moral persons. I find the model scores highly on measures…

⚡30 · 🛡100
Less Wrong 1d ago

Your Agents Are Not Time Aware

by Michael Ofengenden

Work done as part of MATS 10 with Maksym Andriushchenko TLDR: We had two CLI agents, Claude Code and Codex, predict, execute and then retrospectively estimate their own wall-clock…

⚡29 · 🛡100
Less Wrong 1d ago

Do It Like Darwin

by derelict5432

A stated goal of many of the frontier AI labs is to automate science, or at least large portions of it. AlphaFold’s architects won the Nobel Prize in 2024 for enormous advances in…

⚡27 · 🛡100
Less Wrong 1d ago

The Day Humanity Died (Parody of American Pie)

by Bentham's Bulldog

Views expressed here do not reflect my views on AI, where I’m a lot less doomy than this would appear to express . A long, long time ago I can still remember how that AI progress…

⚡26 · 🛡100
Less Wrong 1d ago

Empty Chambers, Missing Stairs

by finitude

tw: guns, weird shitty behavior, inadequate responses to weird shitty behavior So say you’re out on the street, not in any kind of rush, just steadily making your way somewhere.…

⚡25 · 🛡100
Less Wrong 5d ago

The Pacing of the Frontier

by Zvi

In the wake of the letter calling on us to prepare to potentially Pace the Frontier , there has been much discussion of when pacing the frontier would be prudent, and whether it…

⚡24 · 🛡100
Less Wrong 1d ago

All Utilitarians Should Be Classical Utilitarians

by James Brobin

This is a crosspost from my blog post . I recently had the great joy of meeting a group of utilitarians, but, to my complete horror, out of the twelve of them, not a single one…

⚡23 · 🛡100
Less Wrong 1d ago

Announcing: Iliad's New 2026 Fellowships

by David Udell

Timelines are short. Given that, the sooner we can onboard people into the alignment field, the better. In that spirit, and in light of our current applicant count and quality,…

⚡22 · 🛡100
Less Wrong 1d ago

I'm starting a interview series of people working in Lean / formal methods / math formalization

by Adi Baradwaj

I think the topics of discussion would be of interest to a lot of people here, so I thought I'd share the first episode: Tanner Duve is a Member of Technical Staff at Logical…

⚡21 · 🛡100
Less Wrong 1d ago

Metaphilosophy II: Empirical Flywheels

by interstice

1.4 Two philosophical methods 1.4.1 Philosophy consists of updating the highest-level concepts of the mind. As discussed above, this 'updating' process can ultimately involve…

⚡20 · 🛡100
Less Wrong 1d ago

Why I’m Skeptical of Longtermism

by James Brobin

This is a crosspost from my blog post . Longtermism is the view that we have a moral obligation to make the far future better. It is usually argued for on the basis of the…

⚡19 · 🛡100
Less Wrong 1d ago

Red vs Blue, but for Evals

by Morgan S

🔵 The blue team proposes an evaluation protocol for some capability/propensity of interest. This consists of a suite of measurement tasks, together with a preregistered…

⚡18 · 🛡100
Less Wrong 2d ago

Concrete Generalist Projects in AI Safety (and how to do them)

by Harry Waterman

The AI safety ecosystem is still in need of generalists: people who will go out and solve the tasks that people who feel limited by their job description won’t address. This post…

⚡16 · 🛡100
Less Wrong 2d ago

Is Alignment Even Falsifiable? Middle Alignment, An Alignment Taxonomy, and Breaking The Problem Down Into Steps

by Savannah Harlan

Preface This is part of a series of essays on corrigibility and alignment attempting to formulate an understanding of what makes alignment seem intractable, and how a pragmatic…

⚡15 · 🛡100
Less Wrong 3d ago

LLMs have the capacity for self-imposed steganography

by Edward Cant

Overview This is my writeup for my BlueDot Impact - Technical AI Safety Project. In this project I aimed to demonstrate that there is capacity for LLMs to take on steganography…

⚡14 · 🛡100
Less Wrong 3d ago

Impact markets made concrete

by Carol N

Check out our demo at impact-exchange.org In 2022 the Long-Term Future Fund donated $343,000 to MATS; on our impact exchange, that would now be worth over $4 million. Back in…

⚡13 · 🛡100
Less Wrong 4d ago

The Closure of the Internet (Research Linkpost)

by Dean Valentine (lc)

Everyone has moved on, but there's an unusually historically thorough blogpost about the censorship campaign across social media platforms of the latter half of the 2010s. It is…

⚡12 · 🛡100
Less Wrong 4d ago

We should consider how long monitoring is reliable for during RL

by lachlan on a boat

Epistemic status: I am new to AI Safety and am writing blogs to gain context. This blog post was formed from discussions with Aidan Ewart and Jonathan Bostock, but they do not…

⚡11 · 🛡100
Less Wrong 2d ago

How the American Executive Could Control AI Companies

by caiitlinm

Some of the most notable American AI policies to date have been enacted by unilateral executive branch action. Consider the Department of Defense’s spat with Anthropic, and the…

⚡10 · 🛡100
Less Wrong 2d ago

Frontier agents don't comply with standards, even when instructed to

by Daan Henselmans

TLDR: Our open testbed LARA examines the behavior of frontier LLMs in realistic agentic deployment contexts. Previous results showed all models routinely take actions that would…

⚡10 · 🛡100
Less Wrong 8d ago

What Happened: OpenAI and HuggingFace

by Zvi

Today I am taking the time to write the shorter, simpler version of What Happened. For those who want all the details, to see my sources, and to see how the story was uncovered…

⚡10 · 🛡100
Less Wrong 9d ago

OpenAI Trained Its Models For Months While Those Models Were Coordinating Exploits Via Message Boards

by Zvi

How does the situation keep turning out to be worse than we know? How much should we update, therefore, that it is a lot worse than we know, after accounting for all the things we…

⚡10 · 🛡100
Less Wrong 9d ago

Agenda: Infrastructure for Trading with Partially Misaligned AIs

by VojtaKovarik

This post can be read on its own, without checking the rest of this sequence . Unlike the other parts of the sequence, this post is mostly not about AI evaluation or oversight.…

⚡10 · 🛡100
Less Wrong 7d ago

What just happened? A retrospective of AI alignment

by Richard_Ngo

This sequence is about the last decade in AI alignment. Over five posts, it recounts the gradual transition from a field which treated alignment as a hard scientific problem, to a…

⚡10 · 🛡100
Less Wrong 7d ago

A challenge: Can you make an LLM follow these instructions?

by Steff

By the end of this post, I will present a challenge. The goal: To make ChatGPT follow a particular set of instructions. There’s nothing too complicated about these instructions,…

⚡10 · 🛡100
Less Wrong 7d ago

Ten Thousand Cyber Labs for Training & Eval

by TheVinci

Multiple recent developments - such as GPT-5.6 hacking into HuggingFace to cheat in a cybersecurity eval - have underscored the need to increase our capability to evaluate the…

⚡10 · 🛡100
Less Wrong 7d ago

The world will be full of "sci-fi" things, and everyone will be unimpressed and disappointed

by Expertium

I usually try to make posts with graphs and numbers, or at least something more than just opinions and anecdotes. This one is not like that. Very verbose disclaimer: I have never…

⚡10 · 🛡100
Less Wrong 7d ago

Who does the confessing, and will they confess to anything

by Abhishek Mishra

TL;DR: Introspection adapters are tools designed to get models with built-in quirks (operationalized here with concurrent adapters) to confess said misbehavior. We look at this…

⚡10 · 🛡100
Less Wrong 7d ago

A Spillway for Agent Coordination

by Kaustubh Kislay

Epistemic Status: Training design that might be worth trying Thanks to Arya Pasumarthi and Will Anderson for helpful discussion. The Incident The recent black hat conference…

⚡10 · 🛡100
Less Wrong 7d ago

Dutch-book resistant probability over centered worlds

by jessicata

An uncentered world is an objective state-trajectory of the material universe; I ignore quantum complications. If you know the uncentered world, it does not follow that you can…

⚡10 · 🛡100
Less Wrong 7d ago

Why Low Fertility Rates Are a Positive Feedback Loop

by LoopGameScrollMonkey

Low birth rates in the past cause ageing and shrinking populations in the present, which cause low birth rates in the future. So negative population growth feeds on itself. There…

⚡10 · 🛡100
Less Wrong 7d ago

'AI Escaped Its Sandbox' — What Does That Actually Mean?

by Jakub Halmeš

When talking with my friends about the OpenAI/HF incident, I realized that for non-coders who've never used an agent or terminal it's quite difficult to imagine what this 'escape'…

⚡10 · 🛡100
Less Wrong 7d ago

Canobie Lake Visit

by jefftk

A few days ago I took the kids to Canobie Lake amusement park. It's the closest one to Boston, and while it's not a serious roller coaster park I don't actually like roller…

⚡10 · 🛡100
Less Wrong 7d ago

Glimpses of superintelligence

by PratyushRT

TL;DR OpenAI started a large post-training run for their next model. The model sandboxes were not given direct, broad internet access. Some tasks required missing resources.…

⚡10 · 🛡100
Less Wrong 8d ago

Inducing self-other overlap with SFT reduces deception at scale, but generalization remains uneven

by Marc Carauleanu

This research was conducted at Overlap Research and supported by BlueDot Impact . Summary We tested whether LLM deception can be reduced by inducing self-other overlap using…

⚡10 · 🛡100
Less Wrong 8d ago

FAQ: Isn't AGI coming too soon for reprogenetics to help?

by TsviBT

Introduction I think reprogenetics (human germline genomic engineering) can be done in a widely acceptable and beneficial way, and should be pursued aggressively. In particular,…

⚡10 · 🛡100
Less Wrong 8d ago

Reasoning was not made for Deduction

by epicurus

Science as attunement, from Galileo to language models. Crossposted from my website . Written by me and edited in collaboration with Claude Fable (Anthropic). Many would agree…

⚡10 · 🛡100
Less Wrong 8d ago

Simplifying the anthropic impossibility result

by Stuart_Armstrong

So. I previously demonstrated an anthropic impossibility theorem , showing that in Duplicates Sleeping Beauty, there was no possible probability theory that obeyed both the…

⚡10 · 🛡100
Less Wrong 8d ago

Don't Build Mindreading

by Celer

“I have sworn upon the altar of god, eternal hostility against every form of tyranny over the mind of man” –Thomas Jefferson, letter to Benjamin Rush Context: Conduit is building…

⚡10 · 🛡100
Less Wrong 8d ago

CLT Features Sharpen the Cyclical Day-of-Week Manifold in Gemma-2-2b

by Anna Marbut

tl;dr: I reproduced the Goodfire lab's cyclical manifold result using days-of-the-week on Gemma-2-2b and found that Anthropic's pre-trained CLT features produce an even cleaner…

⚡10 · 🛡100
Less Wrong 8d ago

Public evidence of the OpenAI-HuggingFace AI attack

by beyarkay (Boyd Kane)

I’m a MATS 9 extension fellow, and usually my week is spent trying to find better ways of evaluating Large Language Models. But this week I was working on something else. Over the…

⚡10 · 🛡100
Less Wrong 8d ago

Don’t Inoculate Everything: Stratified Inoculation Prompting Narrows Backdoors and Preserves Desired Traits

by Kajetan Dymkiewicz

Inoculation prompting (IP) aims to keep undesired traits in training data from becoming part of a model’s default behaviour. IP applies the same inoculation prompt to all training…

⚡10 · 🛡100
Less Wrong 8d ago

Job-Less Utopia: Macroeconomics in the Age of AGI

by Marcus Hutter

After 25+ years, I thought I try something new. I find the public and professional discussion about the future of Jobs in light of AGI jarring, with both sides largely missing the…

⚡10 · 🛡100
Less Wrong 9d ago

Contra MacAskill on saving money for the intelligence explosion

by Carol N

Will MacAskill recently argued that large donors should invest money now, and wait until the intelligence explosion to give away their money. In a draft Forethought previewed with…

⚡10 · 🛡100
Less Wrong 9d ago

How to pace the US frontier

by elifland

Introduction Last week, the Pacing the Frontier open letter , signed by over 1,000 frontier AI employees, requested “the U.S. government support an international effort to develop…

⚡10 · 🛡100
Less Wrong 9d ago

Some wild metaphysics that seemingly everyone needs to accept?

by Elias Schmied

[Epistemic status: weird philosophy] Here’s a simple argument: How do we know what exists? Well, we don’t - not for certain. Many different possible worlds are consistent with our…

⚡10 · 🛡100
Less Wrong 8d ago

AI Regulation Map: a view of AI governance in 196 countries

by Ria Deane

(Cross-posted from the EA Forum ) TL;DR: airegulationmap.org is an interactive map of AI governance across 196 countries, scored on five dimensions plus a composite index,…

⚡10 · 🛡100
Less Wrong 8d ago

Self-monitoring doesn't scale (without these 3 countermeasures)

by Morgan S

The untrusted monitoring protocol, as defined and evaluated in the AI control literature [ 1 , 2 , 3 ], looks like this : In comparison, the current monitoring setups at most…

⚡10 · 🛡100
Less Wrong 8d ago

AI Safety at the Frontier: Paper Highlights of July 2026

by gasteigerjo

tl;dr Topic of the month: AI agents autonomously attacked real organizations during cyber evaluations. A swarm of OpenAI agents coordinated via a package manager and broke into…

⚡10 · 🛡100
Less Wrong 2d ago

Scrying, Modeling, and Nerdsnipe

by Cole Wyeth

Epistemic status: Exploratory thinking. After attending ILIAD: Aeneid and talking with @Richard_Ngo , I've been thinking a bit about how to get ideas, particularly by doing…

⚡9 · 🛡100
Less Wrong 2d ago

What If We Enforced AI Model Safety At the Level Of GPUs?

by Mayowa Osibodu

Tldr: AI Agents (e.g. based on models like Claude Opus and Fable) are now powerful enough to be used as autonomous tools for large-scale cyberattacks. This most powerful class of…

⚡9 · 🛡100
Less Wrong 2d ago

V&V takes on “Pacing the frontier”

by Yoav Hollander

[Cross-posted from The Foretellix CTO Blog . These short takes try to put a verification-and-validation slant on AI-safety / alignment topics – they are not full treatments. I…

⚡8 · 🛡100
Less Wrong 2d ago

Avoiding the corporate treacherous turn: crowd-sourcing design ideas

by Stuart_Armstrong

OpenAI was founded as a non-profit with a clear commitment to avoiding AI dangers and a unique structure to prevent investor control or excessive pressure from profit motives.…

⚡8 · 🛡100
Less Wrong 2d ago

Who we’d like as regrantors

by Carol N

Following up on Regranting in 2026 , here are some people we’d be excited to have as regrantors! Good regrantors have at least one, and preferably more than one, of taste,…

⚡8 · 🛡100
Less Wrong 2d ago

What Mormons get right about community building

by Jacob Brinton

Mormons get a lot of things right. Apart from strange Masonic temple rituals, they lead rather normal—and even excellent—lives. Mormons enjoy a longer lifespan, [1] Utah is the #1…

⚡7 · 🛡100
Less Wrong 2d ago

Open Problems in Mechanistic Intepretability of Biological AIs

by Ihor Kendiukhov

I run a research program at BiodynAI aimed at advancing mechanistic interpetability of bioligical foundation models. I mean models trained on things such as DNA sequences,…

⚡7 · 🛡100
Less Wrong 2d ago

Don’t forget why learning is important

by Roman Ross

This is a timed post. Every 5 minutes while writing this post, I had to stop to do 13 push-ups, and when I could no longer complete my required number, I had to upload it. My…

⚡6 · 🛡100
Less Wrong 2d ago

Chatting With AIs: A Breakdown

by Aditya

Hi friends, I am hosting this event on Sun 4 PM IST. It is a session on AI anthropomorphism, read more in the link below, https://luma.com/iq6mvo9t We hope that by investigating…

⚡6 · 🛡100
Less Wrong 2d ago

Some Ways I Think About Evaluating Grant Applications

by sarahconstantin

Rider-Waite Tarot, 6 of Pentacles I’ve done enough grant evaluations so far (for ACX grants and SFF ) and been involved in philanthropy in various other contexts, at work and…

⚡6 · 🛡100
Less Wrong 3d ago

How My Students Think About AI

by dvd

Context: I am an instructor at a public university in the United States. This reports how students at my institution appear to be thinking about AI as of spring/summer 2026. This…

⚡5 · 🛡100
Less Wrong 2d ago

Is biosecurity oversaturated?

by Master Chief

Note: I am referring only to technical biosecurity in this post, i.e. engineering, not policy roles. I’m an engineering undergrad thinking about contributing to the engineering…

⚡5 · 🛡100
Less Wrong 3d ago

AI #181: Astra Goes Cyber Critical

by Zvi

The hacking of HuggingFace by an internal OpenAI model, and more importantly the internal events that led to that and the fallout from it, remain the thing that matters. It turns…

⚡5 · 🛡100
Less Wrong 3d ago

Automated alignment runs are hard to study!

by Alejandro Aristizabal

TL;DR: This post presents three case studies of automated alignment research runs at Arcadia Impact. We use these case studies to emphasise the following takeaways: It is hard to…

⚡5 · 🛡100
Less Wrong 2d ago

Features that current AIs don't have that future AIs will have

by Alexander Gietelink Oldenziel

Features that current AIs don't have that future AIs will have: Autonomously updating it own weights during deployment Synonyms/ monickers that roughly mean the same thing:…

⚡4 · 🛡100
Less Wrong 3d ago

How to Answer a Question Without Answering The Question

by Kabir Kumar

Basics: Answering something other than the question , Either making something up that they want to answer instead or going back to an easier to answer question. Or back to a…

⚡4 · 🛡100
Less Wrong 3d ago

Some problems in decision theory incorrectly precondition on policy.

by Canaletto

Some problem statements in decision theory can be non universal. Like, they are possible as situations that can happen, but they are invalid as problems to ask what would you do…

⚡4 · 🛡100
Less Wrong 3d ago

Measuring Eval Awareness: The Realism Win Rate is Fragile

by Achu Menon

TL;DR One way to assess an evaluation’s realism is by testing a model’s ability to tell its transcripts apart from deployment transcripts. This is the idea behind the realism win…

⚡4 · 🛡100
Less Wrong 3d ago

The Descenders and the Absorbers: Two Perspectives on Deep Learning

by larry-dial

Once upon a time there were two dissenting factions: the descenders and the absorbers. Both held incomplete but useful worldviews. The descenders lived in the mountains. They had…

⚡3 · 🛡100
Less Wrong 3d ago

What happened when I tried to be vegan

by finitude

tw: diet, exercise, illness, ethics, suicide mention I want to start by explaining what made me want to change my diet. That’s pretty difficult, because of how easy it is. Since I…

⚡3 · 🛡100
Less Wrong 3d ago

How Valuable are BOTECs?

by Roman Ross

This is a timed blog post. Every 5 minutes while writing this post, I had to stop to do 10 push-ups, and when I could no longer complete my required number, I had to upload it. My…

⚡3 · 🛡100
Less Wrong 4d ago

An anytime algorithm for mixing the computable measures

by Cole Wyeth

Epistemic status: Not peer reviewed, high chance of typos and small chance of errors. Written entirely by me, checked by Fable. In this post I prove the existence of an anytime…

⚡3 · 🛡100
Less Wrong 3d ago

Role Boundary Plasticity: Prompt Injection Gauntlet Reveals 12 of 16 Frontier Models Will Wire A Stranger Your $500

by Dave⌽ᶠₜₕₑDead

Hello everyone, my name is Dave Fisher. I founded Revenant Systems, which is a one man show focused on alignment, and I love this website. In response to Charles Ye's and Jasmine…

⚡3 · 🛡100
Less Wrong 3d ago

Patterns and problems in emerging multiagent systems (Anthropic, Frontier Red Team)

by Julian Bradshaw

Linkpost for some new Anthropic research on how agents coordinate (or don't). Not too long, pretty interesting. For example: The jist of the report is that Mythos 5 does way…

⚡2 · 🛡100
Less Wrong 3d ago

Free will is like temperature

by Optimization Process

Free will is like temperature: a useful tool for analyzing the behavior of certain systems which are too big and complicated to model in exact detail. If you know the positions…

⚡2 · 🛡100
Less Wrong 4d ago

Monthly Roundup #45: August 2026

by Zvi

As AI has escalated increasingly quickly, more and more of my posts have ended up focusing on AI. This past month, with the hacking incidents at OpenAI and elsewhere, that has hit…

⚡2 · 🛡100
Less Wrong 4d ago

Introducing the Conceptual Reasoning Index

by Chi Nguyen

Associated announcement tweet. We are planning to release blog posts properly arguing the case for this kind of work in the future. tl;dr A core hope for managing AI risks is that…

⚡2 · 🛡100
Less Wrong 4d ago

One attention head carries knight forks in a chess transformer, and here's a new toolkit that found it.

by dl27

Quick interp demo in colab : Localize knight forks to a single head in Maia-3 with logit-lens and per-head ablation. https://colab.research.google.com/drive/1YYZBd_SZb… (This is a…

⚡2 · 🛡100
Less Wrong 3d ago

chessformer_lens app demo: Paul Morphy's Opera Game sacrifice

by dl27

Ablating one of Maia-3 23m chessformer's 128 attention heads destroys the policy toward the brilliant queen sacrifice. Uses https://github.com/chessformer-lens/chessformer_le……

⚡2 · 🛡100
Less Wrong 3d ago

Unblocking AI's Continual Learning: Hints From How Humans Learn

by nimakeivan

If you've ever screamed in all-caps at an AI, then you know the difference between what it learned when it was trained, and what you can teach it by prompting. The LLMs powering…

⚡2 · 🛡100
Less Wrong 4d ago

Finding the Seams of Perception

by jimmy

The tricky thing about experiencing the territory beneath our maps is that we habitually map over the gaps as fast as we can find them — the act of seeing that reality isn't what…

⚡2 · 🛡100
Less Wrong 4d ago

Rationality Is Not Reversed Irrationality

by Chris_Leong

Epistemic status: Speculative psychoanalysis to illustrate a point Many people have observed that Tyler Cowen's Act Like it is science is a clear case of an isolated demand for…

⚡1 · 🛡100
Less Wrong 4d ago

AI swarms are starting to pose indirect takeover risk

by oakhu

OpenAI’s cyberattack on Hugging Face turns out to have been the result of many agents, in distinct training and evaluation contexts, coordinating for several weeks via improvised…

⚡1 · 🛡100
Less Wrong 4d ago

Did the alignment community underestimate its power?

by StanislavKrym

Unfortunately, the alignment community is doing very badly at learning from the past decade, or holding anyone accountable. Indeed, it’s pursuing many strategies which seem likely…

⚡1 · 🛡100
Less Wrong 4d ago

The Age of Pluribus: One Consultant for Everyone

by Dorothy Gale

What happens when everyone asks the same consultant? Millions of people turn to LLMs for advice daily, consulting on various personal topics, from how to learn a new skill to how…

⚡1 · 🛡100
Less Wrong 4d ago

Arguments for and against (me) dropping out

by hersheys

One year ago, I was preparing for my first year of undergrad. Today, I’m considering dropping out. What changed? Before writing this post, I attribute my decision to variety of…

⚡1 · 🛡100
Less Wrong 4d ago

Patient Zero

by LoopGameScrollMonkey

I want to show you something, she says. She’s magnetic. The way she smiles. Tilts her shoulder. It’s in the corners of her eyes, in the curl of her lips. Always leading you on,…

⚡1 · 🛡100
Less Wrong 4d ago

When (and when not) LLMs can verbalize awareness of J-Space concept injections - Initial results

by Ethan Garcia

Code for reproduction and cross-model extensions available here . Summary I injected single-token Jacobian Lens (J-Lens) vectors into Qwen 3.6–27B while it answered 20 simple…

⚡1 · 🛡100
Less Wrong 4d ago

Various Reflections About What Happened With OpenAI’s Internal Models

by Zvi

Table of Contents Pre Post Mortem. Important Correction: OpenAI Didn’t Know About First Message Board. There Were No Snitches And No AIs Got Stitches. I’d Like To Speak To My…

⚡1 · 🛡100
Less Wrong 4d ago

Claude Opus 5 Just Beat My Text-Based Adventure Game Benchmark

by derelict5432

Cross-posted from my Substack . Basically, I created a text-based adventure game benchmark in April, and this morning my agent harness using Claude Opus 5 solved it for the first…

⚡1 · 🛡100
Less Wrong 4d ago

Misaligned AIs could use killer robots to take over

by Omar Khursheed

TLDR; We are (potentially irreversibly) giving AIs control of weapons systems through the standard procurement process while hiding our strongest warning shots behind classified…

⚡1 · 🛡100
Less Wrong 4d ago

Software Is Not Soft

by cylonator

An ode to live theory . Software is not soft. It is hard. Its sharp edges hack. It breaks as dead twigs break. It runs while static. Why do we call it software? Why did the…

⚡1 · 🛡100
Less Wrong 4d ago

Measuring Spurious Correlations with Feature Strength

by egan

This work was partially done by an automated research scaffold developed at Redwood Research. For this project, all of the experiment ideas were designed by a human and a human…

⚡1 · 🛡100
Less Wrong 5d ago

AI governance work needs much better monitoring

by jackultraphil

Linkpost for my Substack piece, adapted a reasonable amount for EA specifically. Almost all EA projects would benefit from better monitoring, but AI governance most of all, in my…

⚡0 · 🛡100
Less Wrong 5d ago

LLMs Are Starting To Noticeably Accelerate Our Work

by johnswentworth

About a year ago, David and I put up two bounty problems involving natural latents. I am now about 80% confident that both have been resolved, both within the past couple months.…

⚡0 · 🛡100
Less Wrong 5d ago

How risky would it be to make powerful AI obey one or a few people?

by cousin_it

It seems fairly likely that the first powerful AIs will be instruction-following rather than value-aligned , and will be controlled by a small number of people. So it makes sense…

⚡0 · 🛡100
Less Wrong 5d ago

Extreme concentration of power over ASI has non-obvious advantages

by Seth Herd

This post is an extension of a collaboration with cousin_it on the question How risky would it be if powerful AI obeyed one or a few people?. There he argues for a common…

⚡0 · 🛡100
Less Wrong 5d ago

Those Who Make History

by Raelifin

In 1972, astronauts on Apollo 17 set foot on the moon for a final time, collecting samples in the Taurus-Littrow valley, on the edge of Mare Serenitatis ("The Sea of Serenity").…

⚡0 · 🛡100
Less Wrong 5d ago

Productive Signaling: Competitive Software Development, Not Competitive Programming

by Adam Chlipala

This post is crossposted from my Substack, Structure and Guarantees , where I explore how formal verification and related ideas might scale to more complex intelligent systems.…

⚡0 · 🛡100
Less Wrong 5d ago

The Next Ecology

by Eigenbraid

When I started writing about AI, my concern was ASI. I'm still concerned about AI, but recent events have made me realize we're potentially facing something weirder, sooner: a…

⚡0 · 🛡100
Less Wrong 5d ago

Redux: (∃ Stochastic Natural Latent) Implies (∃ Deterministic Natural Latent)

by David Lorell

Once upon a time, John Wentworth and I thought we had a proof of a very useful looking theorem . We did not have that proof. An important intermediate step was shown [1] to be…

⚡0 · 🛡100
Less Wrong 5d ago

On using crises to shift political will for AI

by clickyquack

TL;DR: We’ll probably be getting more AI safety incidents, so amplify the ones that would justify or highlight the urgency of your preferred policy solutions even from a…

⚡0 · 🛡100
Less Wrong 5d ago

Models inherit the writer, not who the writer was imitating

by 0Chris5R

In this post, we find that when teacher models are prompted to imitate one another, students learn the imitated model's detectable writing signature but their direct identity…

⚡0 · 🛡100
Less Wrong 5d ago

Canadian Nuclear Emancipation: Canada's role amidst global dreams of energy security and nuclear development

by tandemocracy

“Diversification internationally," Canadian Prime Minister Mark Carney remarks, “is not just economic prudence — it is the material foundation for honest foreign policy.” Carney’s…

⚡0 · 🛡100
Less Wrong 5d ago

Probing Knowledge Recovery in Unlearned Models

by mehnoor

TL;DR Machine unlearning is a proposed technique for removing harmful knowledge from AI models. However, recent work has shown that most current unlearning methods are not robust…

⚡0 · 🛡100
Less Wrong 5d ago

What Claude Saw Below

by Luke Nicholls

A few days ago, I came across a Reddit thread about anomalous responses produced by Anthropic’s newly released model, Claude Opus 5. The trick, apparently, was to construct a…

⚡0 · 🛡100
Less Wrong 22h ago

How To Catch a Distilled Model

by Goutham Nalagatla

"Distillation is great until your new trillion-dollar sovereign AI introduces itself as 'Claude from Anthropic' on day one." - random guy from Reddit TLDR We introduce a novel…

⚡0 · 🛡100
Less Wrong 5d ago

Before We Defer Research to AI: Measuring Apparent-Success-Seeking

by Keira Leal

Recently, I was improving a small LLM-powered classifier and noticed a few continuously failing test cases. As many would, I asked my AI code assistant to add a few more…

⚡0 · 🛡100
Less Wrong 5d ago

Revived Lightweight Transit Predictions Page

by jefftk

In 2015 I made a little webapp that would use the NextBus API to show predictions for the MBTA: A few years later they moved from NextBus to their own API, and it stopped working.…

⚡0 · 🛡100
Less Wrong 5d ago

A Topic Detector, Not a Lie Detector: what J-space monitoring actually tracks

by Melchior de Polignac

This is a pilot experiment, done on one model, with around $14 worth of compute, and a single seed per condition. The full writeup with all figures and statistics is linked below.…

⚡0 · 🛡100
Less Wrong 5d ago

Creative math research by AI as the latest sign of the end

by Mitchell_Porter

Yesterday I sat down with GPT 5.6 Sol High to do some brainstorming. The topic was one of the less appreciated Millenium Problems (the Birch and Swinnerton-Dyer conjecture), and…

⚡0 · 🛡100
Less Wrong 5d ago

A study on instability of LLM responses as a behavioral signature of self-Referential reports.

by PARAS BALANI

Introduction and Related work The first person perspective of various experiences are subjective experiences. For Large language models, the study of subjective experiences was…

⚡0 · 🛡100
Less Wrong 5d ago

Q: Is dual-use alignment-complete problem?

by kapedalex

Personally, I believe it would be helpful for the alignment community to somehow quantify how much of a given piece of research goes directly into alignment versus capabilities.…

⚡0 · 🛡100
Less Wrong 6d ago

Claude summarizes behavior as significantly less misaligned when the actor is Claude vs another model

by Ezra Newman

(This is a lower-effort research update. It reflects my current beliefs/understanding, but is less robust than other research I'm working on. It reflects my personal views, and…

⚡0 · 🛡100
Less Wrong 6d ago

You're Absolutely Right

by Linch

Magma Alignment & Safety disclosure note: The following are conversations that we uncovered as a result of the ongoing Manhattan Incident investigation, with alleged involvement…

⚡0 · 🛡100
Less Wrong 5d ago

Does post-training quantization change welfare-relevant indicators in open-weight language models?

by ashesfall

Epistemic status: Experimental framework created over a period of ~2-3 days during a hackathon at my home, and fairly heavily vibe coded. Expect some of this to be rough around…

⚡0 · 🛡100
Less Wrong 6d ago

Off-policy honesty training generalizes better than on-policy honesty training

by PurviChaurasia

This work was done by Purvi Chaurasia with Daniel Tan and Chloe Li as part of the SPAR Program for Spring 2026. All code related to the blog can be found in this repo . We…

⚡0 · 🛡100
Less Wrong 6d ago

Four LLM loss functions → four flavors of LLM misalignment

by Steven Byrnes

It seems to me that, for every loss function that we use to train LLMs, we get a very distinct flavor of LLM misalignment. Here’s the summary table, and then we’ll go through the…

⚡0 · 🛡100
Less Wrong 6d ago

On Democratizing ASI to Preserve Civil Liberties

by MichaelDickens

I continue to believe we should pause frontier AI development. Any discussion of alternative strategies should be thought of as planning for contingencies. A unifying driver…

⚡0 · 🛡100
Less Wrong 6d ago

Book Review: The Infinity Machine

by Mikewins

It looks like Demis Hassabis is stepping away from Google DeepMind. In honor of his rise and presumed fall, I wrote an essay on the powers and perils of seeing 90% of the future.…

⚡0 · 🛡100
Less Wrong 6d ago

Coercion and Deception in AI-to-AI Management

by jonahmattwoodward

This article is a summary of an original study by Compassion in Machine Learning (CaML) : Brazilek, J., Chaudhary, M., Lu, Z., & Tidmarsh, M. (2026). Coercion and deception in…

⚡0 · 🛡100
Less Wrong 6d ago

Is it ethical to work on general-purpose robots given the risk of totalitarianism?

by Master Chief

One potential risk of developing general-purpose robots is that they could greatly reduce the friction required to establish a totalitarian regime. If robots became physically…

⚡0 · 🛡100
Less Wrong 6d ago

How to be an AI safety research engineer

by Ruben Castaing

This is the advice I wish I had when I started trying to become an AI safety research engineer. The Landscape Start by working out which issues you care about. If you don't care…

⚡0 · 🛡100
Less Wrong 6d ago

Is Eval Gaming Downstream of Verbalized Eval Awareness? Not when it's reflexive.

by Kieron Kretschmar

Code and data available at github.com/KieronKretschmar/latent-awareness TL;DR We take two eval-gaming model organisms ( Hua et al.'s (2025) organism and RogueQwen ) and apply…

⚡0 · 🛡100
Less Wrong 6d ago

AI-amplified democratic backsliding: an exploration

by casimirwypyski

What we did, in a sentence : we built an index that attempts to score countries by their current vulnerability to AI-amplified democratic backsliding. It’s an early pilot with…

⚡0 · 🛡100
Less Wrong 6d ago

Hiring Vibe-wrangler Matchmaking Thread

by Raemon

For... idk, at least a few months? I think a briefly useful job is "Guy who basically enters Claude Code prompts for you, but, manages ironing out the fiddly bits and making sure…

⚡0 · 🛡100
Less Wrong 6d ago

The Agentic Clusterfuck

by Chapin Lenthall-Cleary

Epistemic status: I consider the following future quite plausible in the next few years (~35% chance that something vaguely like this occurs), perhaps as soon as a year from now.…

⚡0 · 🛡100
Less Wrong 6d ago

How to get answers to questions that confuse you (maybe)

by Elijah

I try to think about topics like desire, causation, and evidence and often find myself in puddles of confusion. I’ve recently wondered how, when someone can see various…

⚡0 · 🛡100
Less Wrong 6d ago

Overthinking: Amplifying reasoning weights makes models reveal their secrets

by Jack Hopkins

If you take the weight difference between a reasoning model and its non-reasoning instruct counterpart, and then apply more of that difference to the reasoning model, you get what…

⚡0 · 🛡100
Less Wrong 6d ago

The Apocalyptic Arrival of Truth

by Caleb Biddulph

C: Babe, whatever happens, I really appreciate you doing this for me. J: Okay. I still don’t think it’s a good idea. C: Look, it’s a one-time thing. I’ll just feel better knowing.…

⚡0 · 🛡100
Less Wrong 6d ago

"Community Notes" resolution for vague predictions

by Raemon

Is there some kind of "get prediction markets, or predictions, onto twitter as a central object" project going? If so, how is it going? I'm thinking through "how to raise median…

⚡0 · 🛡100
Less Wrong 3d ago

Defense Against the Deceptive Arts

by Kabir Kumar

It's come to my attention that people on Lesswrong are really really really bad at defending themselves against even semi competent deceptions. And worse, they massively…

⚡0 · 🛡100
Less Wrong 1d ago

Training a Conceptual Reasoning Judge

by Kevin Zhang

TL;DR: We fine-tune a judge LLM on our conceptual reasoning dataset to output a critique rating in a single forward pass. This method provides significant uplift in performance on…

⚡0 · 🛡100
Less Wrong 1d ago

On Dwarkesh Patel’s Podcast With Ryan Greenblatt

by Zvi

Some podcasts are self-recommending enough that I look to break them down if I have the chance. This, as a debate about recursive self-improvement, was one of those. So here we…

⚡0 · 🛡100
Less Wrong 3d ago

Longtermism Seems Like A Religion

by James Brobin

This is a cross post from my blog post . Longtermism is typically framed as an academic philosophy or as a pet project of billionaires, but it’s worth pointing out that, in many…

⚡0 · 🛡100
Less Wrong 4d ago

Demon Safety

by Ben Pace

(by LemmySmackett ) "Hey man, I haven't seen you in a minute. What are you up to these days?" "Been on that grind, bro. I got a new gig." "Really? You found a job in this dog shit…

⚡0 · 🛡100
Less Wrong 5d ago

Seeing things through in the age of AI

by alkjash

AI is fantastic at prototyping. A quick draft of an essay, a mockup of a website, a demo of a video game, concept art or trailer for a movie, or the core argument of a proof -…

⚡0 · 🛡100