AI 投研
AI 赛道深度、公司拆解、概念解读和周报月报。
Early Indicators of Reward Hacking via Reasoning Interpolation
AI 摘要:Using importance sampling with fine-tuned donor prefil...
OpenAI hit with multistate probe into possible user harm as its IPO looms
AI 摘要:OpenAI received a subpoena from several states as part...
SAEs trained on the same data don’t learn the same features
AI 摘要:In this post, we show that when two TopK SAEs are trai...
Mechanistic Anomaly Detection Research Update 2
AI 摘要:Interim report on ongoing work on mechanistic anomaly ...
Third-party evaluation to identify risks in LLMs’ training data
AI 摘要:An overview of the minetester and preliminary work
Partially rewriting an LLM in natural language
AI 摘要:Using interpretations of SAE latents to simulate activ...
VINC-S: Closed-form Optionally-supervised Knowledge Elicitation with Paraphrase Invariance
本文来自EleutherAI官方博客,介绍了基于2023年春季项目成果的VINC-S方法,这是一种具备释义不变性的闭式可...
Experiments in Weak-to-Strong Generalization
AI 摘要:Writing up results from a recent project
Free Form Least-Squares Concept Erasure Without Oracle Concept Labels
AI 摘要:Achieving even more surgical edits than LEACE without ...
Mechanistic Anomaly Detection Research Update
AI 摘要:Interim report on ongoing work on mechanistic anomaly ...