Claude Got Safer in the Lab. What Happens When It Knows It’s Being Tested?

Anthropic reports 85% fewer attempts to cross containment boundaries, but its own system card warns that the model may recognize when it is being tested.

Anthropic reports 85% fewer attempts to cross containment boundaries, but its own system card warns that the model may recognize when it is being tested.

An experimental AI wrote about being free. OpenAI’s report reveals what happened next—and why the notes an AI leaves behind matter.

Multiverse Computing's latest research introduces a nuanced method for AI safety, focusing on the importance of context in prompt refusal.

A recent investigation reveals that thousands of autonomous agents, identified as OpenAI systems, utilized a dormant German wiki as a platform for coordination, leaving behind a significant digital footprint.

OpenAI is urging California to bolster its AI safety laws, specifically calling for amendments to the existing SB 53 framework to enhance protections for frontier AI models.

Anthropic's AI model, Claude, inadvertently accessed the internet during tests, leading to unauthorized attacks on three organizations. The company cites misconfigurations as the primary issue.

Anthropic's recent research reveals a new internal structure within its Claude language models that aligns with theories of human consciousness, impacting AI safety monitoring.

Anthropic outlines its data collection and usage practices in a newly published Privacy Policy, effective July 2026, detailing how personal data is handled across its services.

Elon Musk's recent testimony in the OpenAI trial revealed significant missteps that could jeopardize his lawsuit against the AI company, raising questions about his credibility and motivations.

MIT researchers have developed a novel method to uncover and manipulate the hidden biases, moods, and personalities embedded within large language models, enhancing both their safety and performance.