Inside OpenAI’s First Chip

The Data Exchange ·

Richard Ho describes OpenAI’s chip work as an effort to lower inference and compute costs. He explains the memory-compute locality built into the architecture and how one device can operate along the Pareto curve between high throughput and low latency. Lisez 3 points de vue avec leurs éléments à l’appui et les liens vers les sources.

Ben LoricaRichard Ho

En un coup d’œil

Moments clés3

Extraits courts et attribués, avec le contexte permettant de les vérifier. La conversation complète reste chez son éditeur.

AI infrastructure strategy

Mission: lower inference cost

Extrait original

They wanted to look really hard at what we could do to lower the cost of inference and lower the cost of compute.
Contexte

I think this project really kicked off when Sam and Greg decided that infrastructure was going to be a key differentiator and one of the key drivers of AI and intelligence to users.

Hardware flexibility

Single-device Pareto curve for throughput and latency

Extrait original

What we have in a single device—and I think this is really critical for our infrastructure and being able to lower the cost of the infrastructure—is a single device that can operate at both points, or at any point along that curve. We call it the Pareto curve between high throughput and low latency.

Source et méthodologie

Ces points de vue renvoient à leurs sources originales. Les reformulations sont signalées et ne sont pas des citations mot à mot.

Ouvrir la transcription ou les documents sources (s’ouvre dans un nouvel onglet)Signaler un problème