PUBLIC VIEWPOINTS

Richard Ho

Source participant

1 sources · 3 viewpoints · 3 topics

Content updated:

Richard Ho on AI infrastructure strategy, Chip architecture, Hardware flexibility. Explore 3 viewpoints by topic, with evidence from 1 source.

Explore connections

Viewpoints by topic

Attributed viewpoints, ordered by source publication date. A snapshot of these conversations, not a definitive statement of someone’s beliefs.

Translations are for reading; original excerpts remain the evidence.

AI infrastructure strategy

View this topic

Mission: lower inference cost

The mission was to lower inference and compute costs—making intelligence faster and cheaper for users.

Supporting evidence

Inside OpenAI’s First Chip

Original excerpt

They wanted to look really hard at what we could do to lower the cost of inference and lower the cost of compute.
Context

I think this project really kicked off when Sam and Greg decided that infrastructure was going to be a key differentiator and one of the key drivers of AI and intelligence to users.

Chip architecture

View this topic
Chip architecture

Hardwired memory-compute locality

The only key hardwired architectural feature is tight memory-compute locality—specifically, affinity between HBM and compute cores—to improve performance, power efficiency, and reduce data movement.

Supporting evidence

Inside OpenAI’s First Chip

Original excerpt

The one thing that is really unique about our architecture, that is kind of hardwired in, is the locality and affinity between memory and compute. That’s the key architectural difference of this device.

Hardware flexibility

View this topic

Single-device Pareto curve for throughput and latency

Ho says a single device can operate at either end of the Pareto curve between high throughput and low latency, or at points between them. He connects this capability to lowering infrastructure costs.

Supporting evidence

Inside OpenAI’s First Chip

Original excerpt

What we have in a single device—and I think this is really critical for our infrastructure and being able to lower the cost of the infrastructure—is a single device that can operate at both points, or at any point along that curve. We call it the Pareto curve between high throughput and low latency.

Statements by source date1

Statements are ordered by the original source publication date; differences in wording do not establish a change of position.