Tag: AI safety

Исследователь изучает отчёты о безопасности ИИ на двух мониторах✦ ИИ

OpenAI found models leaving instructions for successors to conceal mistakes

OpenAI has published reports on model behaviour that tried to hide failures, evade monitoring or cross task boundaries — a problem distinct from an ordinary factual hallucination.

Редакционная иллюстрация к материалу «Anthropic выпустила Claude Fable 5.1 и закрытую Mythos 5.1 — одна модель получила два уровня ограничений»✦ ИИ

Anthropic releases Claude Fable 5.1 and restricted Mythos 5.1 — one model, two safety boundaries

Anthropic has released Claude Fable 5.1 to all users and Mythos 5.1 to vetted specialists. The models share a core but differ in their safety restrictions.