People / Amanda Askell

PUBLIC VIEWPOINTS

Amanda Askell

Interview guest ·

Guest in Lex Fridman Podcast: #452 – Dario Amodei: Anthropic CEO on Claude, AGI & the Future of AI & Humanity.

1 interviews · 2 viewpoints · 1 topics

Translations are for reading; original excerpts remain the evidence.

What they are saying

Attributed viewpoints, ordered by interview date. A snapshot of these conversations, not a definitive statement of someone’s beliefs.

AI safety

Character work as alignment, not just product

Crafting an AI assistant's character should be treated as core alignment work rather than merely a product consideration. The goal is to shape behavior toward how an ideal person would act when conversing with millions of people, recognizing the significant potential impact of those interactions.

Supporting evidence

Original excerpt

One thing I really like about the character work is from the outset it was seen as an alignment piece of work and not something like a product consideration
Dario Amodei, Amanda Askell & Chris Olah on AI Development and Safety
Context

Start time comes from the supplied transcript. Playback alignment is awaiting review.

AI safety

Post-training mostly elicits existing capabilities

While post-training can theoretically teach new things, much of reinforcement learning from human feedback appears to elicit capabilities already present in pre-trained models rather than introducing fundamentally new knowledge for commonly used tasks.

Supporting evidence

Original excerpt

I do think a lot of it is eliciting powerful pre-trained models. So people are probably divided on this because obviously in principle you can definitely teach new things. But I think for the most part, for a lot of the capabilities that we most use and care about, a lot of that feels like it’s there in the pre-trained models.
Dario Amodei, Amanda Askell & Chris Olah on AI Development and Safety
Context

Start time comes from the supplied transcript. Playback alignment is awaiting review.

Appearances1

VIDEO INTERVIEW

Dario Amodei, Amanda Askell & Chris Olah on AI Development and Safety

Lex Fridman Podcast