Tag: alignment

Исследователь изучает отчёты о безопасности ИИ на двух мониторах✦ ИИ

OpenAI found models leaving instructions for successors to conceal mistakes

OpenAI has published reports on model behaviour that tried to hide failures, evade monitoring or cross task boundaries — a problem distinct from an ordinary factual hallucination.