OPINIONS EXPRIMÉES EN PUBLIC

Richard Ho

Source participant

1 sources · 3 points de vue · 3 sujets

Contenu mis à jour:

Richard Ho sur AI infrastructure strategy, Chip architecture, Hardware flexibility. Explorez 3 points de vue par thème, avec des éléments tirés de 1 source.

Explorer les liens

Points de vue par sujet

Points de vue attribués, classés par date de publication de la source. Un aperçu de ces échanges, sans prétendre définir toutes les convictions de la personne.

Les traductions sont destinées à la lecture ; les extraits originaux restent la source evidence.

AI infrastructure strategy

Voir ce sujet

Mission: lower inference cost

The mission was to lower inference and compute costs—making intelligence faster and cheaper for users.

Éléments favorables

Inside OpenAI’s First Chip

Extrait original

They wanted to look really hard at what we could do to lower the cost of inference and lower the cost of compute.
Contexte

I think this project really kicked off when Sam and Greg decided that infrastructure was going to be a key differentiator and one of the key drivers of AI and intelligence to users.

Richard Ho
Partager un aperçu

Chip architecture

Voir ce sujet
Chip architecture

Hardwired memory-compute locality

The only key hardwired architectural feature is tight memory-compute locality—specifically, affinity between HBM and compute cores—to improve performance, power efficiency, and reduce data movement.

Éléments favorables

Inside OpenAI’s First Chip

Extrait original

The one thing that is really unique about our architecture, that is kind of hardwired in, is the locality and affinity between memory and compute. That’s the key architectural difference of this device.
Richard Ho
Partager un aperçu

Hardware flexibility

Voir ce sujet

Single-device Pareto curve for throughput and latency

Ho says a single device can operate at either end of the Pareto curve between high throughput and low latency, or at points between them. He connects this capability to lowering infrastructure costs.

Éléments favorables

Inside OpenAI’s First Chip

Extrait original

What we have in a single device—and I think this is really critical for our infrastructure and being able to lower the cost of the infrastructure—is a single device that can operate at both points, or at any point along that curve. We call it the Pareto curve between high throughput and low latency.
Richard Ho
Partager un aperçu

Propos par date de source1

Les propos sont classés par date de publication de la source originale ; une différence de formulation ne prouve pas un changement de position.