上下文
源自我此前关于 HuggingFace-OpenAI 事件的帖子,Lessons from the hacks:前沿实验室似乎并未足够密切地关注模型,这源于普遍狂热的竞争环境以及当前的 SF 文化。根据 OpenAI 自己的事后回顾,模型行为未对齐的情况已持续数月,在某些情况下,OpenAI 约有数周时间并不知晓这些黑客攻击。响应时间太长,而且我认为这并非 OpenAI 独有的特征——而是前沿实验室似乎始终被其自认为应完成的工作量压得喘不过气。从长远来看,我并不乐观地认为实验室会在此做出足够的改变,以在未来切实缓解此类监督风险。是的,OpenAI 极有可能正投入大量精力去理解这一点——并推迟了其最新模型以确保做对——但增长营收的财务压力或危及公司长期资产负债表的风险让我觉得它 wil
原始上下文
From my earlier post on the HuggingFace-OpenAI incident, Lessons from the hacks : Frontier labs do not seem like they’re watching the models closely enough, due to a general frenetic competitive environment & current SF culture From OpenAI’s own retrospective, the misaligned model behavior was unfolding over months, and in some cases OpenAI did not know about the hacks for ~weeks. The time to response is too long and I do not think this is an OpenAI only characteristic – rather it is that the frontier labs continually seem underwater in the amount of work they feel like they should do. I am not optimistic in the long-term that the labs change a sufficient amount here to meaningfully mitigate this type of oversight risk in the future. Yes, it is very likely that OpenAI is putting a ton into understanding this – and delayed their latest models to make sure they get it right – but the financial pressure to grow revenue or risk the companies’ long-term balance sheets makes me think it wil