First Rival Android App Store Arrives In the US Play Store
Aptoide has become the first third-party Android app store available directly through Google Play in the U.S. "The change is a direct result of Google's litigation with Epic,…
A Topic Detector, Not a Lie Detector: what J-space monitoring actually tracks
This is a pilot experiment, done on one model, with around $14 worth of compute, and a single seed per condition. The full writeup with all figures and statistics is linked below.…
How To Catch a Distilled Model
"Distillation is great until your new trillion-dollar sovereign AI introduces itself as 'Claude from Anthropic' on day one." - random guy from Reddit TLDR We introduce a novel…
Creative math research by AI as the latest sign of the end
Yesterday I sat down with GPT 5.6 Sol High to do some brainstorming. The topic was one of the less appreciated Millenium Problems (the Birch and Swinnerton-Dyer conjecture), and…
A study on instability of LLM responses as a behavioral signature of self-Referential reports.
Introduction and Related work The first person perspective of various experiences are subjective experiences. For Large language models, the study of subjective experiences was…
Q: Is dual-use alignment-complete problem?
Personally, I believe it would be helpful for the alignment community to somehow quantify how much of a given piece of research goes directly into alignment versus capabilities.…
You're Absolutely Right
Magma Alignment & Safety disclosure note: The following are conversations that we uncovered as a result of the ongoing Manhattan Incident investigation, with alleged involvement…
Claude summarizes behavior as significantly less misaligned when the actor is Claude vs another model
(This is a lower-effort research update. It reflects my current beliefs/understanding, but is less robust than other research I'm working on. It reflects my personal views, and…
Does post-training quantization change welfare-relevant indicators in open-weight language models?
Epistemic status: Experimental framework created over a period of ~2-3 days during a hackathon at my home, and fairly heavily vibe coded. Expect some of this to be rough around…
Off-policy honesty training generalizes better than on-policy honesty training
This work was done by Purvi Chaurasia with Daniel Tan and Chloe Li as part of the SPAR Program for Spring 2026. All code related to the blog can be found in this repo . We…
A Zoom Screen-Sharing Bug Let Anyone Take Over Other Devices On a Call
An anonymous reader quotes a report from Wired: As AI models gain advanced capabilities to find vulnerabilities in software, develop ways to exploit them, and even carry out…
Four LLM loss functions → four flavors of LLM misalignment
It seems to me that, for every loss function that we use to train LLMs, we get a very distinct flavor of LLM misalignment. Here’s the summary table, and then we’ll go through the…
On Democratizing ASI to Preserve Civil Liberties
I continue to believe we should pause frontier AI development. Any discussion of alternative strategies should be thought of as planning for contingencies. A unifying driver…
Disneyland Announces Star Wars/Fortnite Collaboration, 'Avatar' Attraction, and a Newer Tomorrowland
Saturday Disneyland announced a special Star Wars-themed collaboration with third-person shooter game Fortnite — including a slick new trailer for its Star Wars: Smuggler's Gambit…
Think Tank Battling Progressive Policies Gets Funding From Insurance Industry
Third Way has announced it is “preparing for the next war” within the Democratic Party against democratic socialism.
Trump Blames Iran for Gaza Genocide as He Demands “Compensation” From Tehran
The post came in response to Iranian officials demanding that the US compensate Iran for the war.