Topics / AI safety

A TOPIC, IN CONTEXT

AI safety

Model limitations, evidence grounding and responsible deployment.

5 people · 3 sources · 8 viewpoints

Viewpoint map

Explore by person. Select two or three to compare.

5 people · 3 sources · 8 viewpoints

Jensen Huang

Giving agentic systems only two of three capabilities at once

Jensen Huang describes allowing agentic systems only two of three capabilities at a time: accessing sensitive information, executing code and communicating externally. He also describes applying enterprise access controls and connecting the systems to a policy engine.

Supporting evidence

Original excerpt

We give you two out of three rights. Agentic systems can access sensitive information, it can execute code, and it can communicate externally. We could keep things safe if we gave you two out of those three capabilities at any time, but not all three.
Jensen Huang on AI agent safety, leadership and trust
Context

They install super easy. It makes sure that it’s secure. So you eloquently explained how we have a long history of blockers that we thought were going to be blockers, and we overcame them. But now looking into the future, what do you think might be the blockers now that it’s clear that agents will be everywhere? So it’s obviously we’re going to need compute. So what is going to be the blocker for that scaling?

Start time comes from the supplied transcript. Playback alignment is awaiting review.

Share insightCheck this claim

Dario Amodei

Internet data quality constraints

Amodei acknowledges that running out of high-quality internet data is a possible scaling limit. While trillions of words exist online, much is repetitive, low-quality search engine optimization content, or potentially future AI-generated text, constraining useful training material.

Supporting evidence

Original excerpt

we simply run out of data. There’s only so much data on the internet, and there’s issues with the quality of the data. You can get hundreds of trillions of words on the internet, but a lot of it is repetitive or it’s search engine optimization drivel
Dario Amodei, Amanda Askell & Chris Olah on AI Development and Safety
Context

Start time comes from the supplied transcript. Playback alignment is awaiting review.

Share insightCheck this claim

Current computer-use models require boundaries and guardrails

The current computer-use model still makes mistakes and misclicks, so users cannot leave it running unattended for extended periods. Anthropic released it first as an API rather than giving consumers direct computer control, emphasizing the need for boundaries and guardrails as capabilities improve.

Supporting evidence

Original excerpt

It makes mistakes, it misclicks. We were careful to warn people, “Hey, you can’t just leave this thing to run on your computer for minutes and minutes. You got to give this thing boundaries and guardrails.” And I think that’s one of the reasons we released it first in an API form
Dario Amodei, Amanda Askell & Chris Olah on AI Development and Safety
Context

Start time comes from the supplied transcript. Playback alignment is awaiting review.

Share insightCheck this claim

Poorly targeted AI regulation risks creating backlash against safety

If AI regulation is poorly targeted and creates unnecessary burdens, practitioners may conclude safety concerns are exaggerated after wasting time on compliance. Badly designed regulation could generate durable consensus against future accountability measures, making surgical targeting essential.

Supporting evidence

Original excerpt

if we get something in place that’s poorly targeted, that wastes a bunch of people’s time, what’s going to happen is people are going to say, “See, these safety risks, this is nonsense. I just had to hire 10 lawyers to fill out all these forms.
Dario Amodei, Amanda Askell & Chris Olah on AI Development and Safety
Context

Start time comes from the supplied transcript. Playback alignment is awaiting review.

Share insightCheck this claim

Amanda Askell

Character work as alignment, not just product

Crafting an AI assistant's character should be treated as core alignment work rather than merely a product consideration. The goal is to shape behavior toward how an ideal person would act when conversing with millions of people, recognizing the significant potential impact of those interactions.

Supporting evidence

Original excerpt

One thing I really like about the character work is from the outset it was seen as an alignment piece of work and not something like a product consideration
Dario Amodei, Amanda Askell & Chris Olah on AI Development and Safety
Context

Start time comes from the supplied transcript. Playback alignment is awaiting review.

Share insightCheck this claim

Post-training mostly elicits existing capabilities

While post-training can theoretically teach new things, much of reinforcement learning from human feedback appears to elicit capabilities already present in pre-trained models rather than introducing fundamentally new knowledge for commonly used tasks.

Supporting evidence

Original excerpt

I do think a lot of it is eliciting powerful pre-trained models. So people are probably divided on this because obviously in principle you can definitely teach new things. But I think for the most part, for a lot of the capabilities that we most use and care about, a lot of that feels like it’s there in the pre-trained models.
Dario Amodei, Amanda Askell & Chris Olah on AI Development and Safety
Context

Start time comes from the supplied transcript. Playback alignment is awaiting review.

Share insightCheck this claim

Chris Olah

Mechanistic interpretability reveals deception-related features

Researchers found multiple features related to deception, lying, withholding information, and power-seeking in neural networks. Forcing these features active causes the model to exhibit corresponding undesirable behaviors, though this area of research remains in early stages.

Supporting evidence

Original excerpt

we find quite a few features related to deception and lying. There’s one feature where it fires for people lying and being deceptive, and you force it active and Claude starts lying to you.
Dario Amodei, Amanda Askell & Chris Olah on AI Development and Safety
Context

Start time comes from the supplied transcript. Playback alignment is awaiting review.

Share insightCheck this claim

Pieter Levels

EU AI regulation favors incumbents over startups

Levels argues that heavy European regulation creates regulatory capture, benefiting large incumbent companies that can afford compliance while blocking newcomers. He contrasts this with America, where he claims he can start an AI startup immediately by just opening his laptop.

Supporting evidence

Original excerpt

regulation is very good for big companies because they can follow it. I can’t follow it, right? If I want to start an AI startup in Europe now, I cannot because there’s an AI regulation that makes it very complicated for me. I probably need to get notaries involved.
Pieter Levels on Bootstrapping Startups, AI Image Models, and Indie Development
Context

Start time comes from the supplied transcript. Playback alignment is awaiting review.

Share insightCheck this claim

These are individual perspectives, not a measure of consensus. Source material stays in its original language.