OPINIONES PÚBLICAS

Richard Ho

Source participant

1 fuentes · 3 opiniones · 3 temas

Contenido actualizado:

Richard Ho sobre AI infrastructure strategy, Chip architecture, Hardware flexibility. Explora 3 puntos de vista por tema, con evidencias de 1 fuente.

Explorar conexiones

Perspectivas por tema

Opiniones atribuidas, ordenadas por fecha de publicación de la fuente. Una muestra de estas intervenciones, no una definición completa de las creencias de la persona.

Las traducciones son para facilitar la lectura; los extractos originales siguen siendo la source evidence.

AI infrastructure strategy

Ver este tema

Mission: lower inference cost

The mission was to lower inference and compute costs—making intelligence faster and cheaper for users.

Evidencia a favor

Inside OpenAI’s First Chip

Extracto original

They wanted to look really hard at what we could do to lower the cost of inference and lower the cost of compute.
Contexto

I think this project really kicked off when Sam and Greg decided that infrastructure was going to be a key differentiator and one of the key drivers of AI and intelligence to users.

Richard Ho
Compartir información

Chip architecture

Ver este tema
Chip architecture

Hardwired memory-compute locality

The only key hardwired architectural feature is tight memory-compute locality—specifically, affinity between HBM and compute cores—to improve performance, power efficiency, and reduce data movement.

Evidencia a favor

Inside OpenAI’s First Chip

Extracto original

The one thing that is really unique about our architecture, that is kind of hardwired in, is the locality and affinity between memory and compute. That’s the key architectural difference of this device.
Richard Ho
Compartir información

Hardware flexibility

Ver este tema

Single-device Pareto curve for throughput and latency

Ho says a single device can operate at either end of the Pareto curve between high throughput and low latency, or at points between them. He connects this capability to lowering infrastructure costs.

Evidencia a favor

Inside OpenAI’s First Chip

Extracto original

What we have in a single device—and I think this is really critical for our infrastructure and being able to lower the cost of the infrastructure—is a single device that can operate at both points, or at any point along that curve. We call it the Pareto curve between high throughput and low latency.
Richard Ho
Compartir información

Declaraciones por fecha de la fuente1

Las declaraciones se ordenan por fecha de publicación de la fuente original; las diferencias de redacción no demuestran un cambio de postura.