Early Indicators of Reward Hacking via Reasoning Interpolation

AI 摘要:Using importance sampling with fine-tuned donor prefil...

ybx-ai-radar
2026-06-15 17:09
OpenAI hit with multistate probe into possible user harm as its IPO looms

AI 摘要:OpenAI received a subpoena from several states as part...

ybx-ai-radar
2026-06-15 16:37
SAEs trained on the same data don’t learn the same features

AI 摘要:In this post, we show that when two TopK SAEs are trai...

ybx-ai-radar
2026-06-15 16:05
Mechanistic Anomaly Detection Research Update 2

AI 摘要:Interim report on ongoing work on mechanistic anomaly ...

ybx-ai-radar
2026-06-15 15:27
Third-party evaluation to identify risks in LLMs’ training data

AI 摘要:An overview of the minetester and preliminary work

ybx-ai-radar
2026-06-15 14:53
Partially rewriting an LLM in natural language

AI 摘要:Using interpretations of SAE latents to simulate activ...

ybx-ai-radar
2026-06-15 14:20
VINC-S: Closed-form Optionally-supervised Knowledge Elicitation with Paraphrase Invariance

本文来自EleutherAI官方博客,介绍了基于2023年春季项目成果的VINC-S方法,这是一种具备释义不变性的闭式可...

ybx-ai-radar
2026-06-15 13:47
Experiments in Weak-to-Strong Generalization

AI 摘要:Writing up results from a recent project

ybx-ai-radar
2026-06-15 13:11
Free Form Least-Squares Concept Erasure Without Oracle Concept Labels

AI 摘要:Achieving even more surgical edits than LEACE without ...

ybx-ai-radar
2026-06-15 12:38
Mechanistic Anomaly Detection Research Update

AI 摘要:Interim report on ongoing work on mechanistic anomaly ...

ybx-ai-radar
2026-06-15 12:06