Why do standard GPUs struggle with data movement during the decode phase of inference?

Andrew Feldman
Replied byAndrew Feldman

Co-founder & CEO at Cerebras Systems

Niche: Technology
Revenue: Approx. $72 Million/month
Location: Sunnyvale, California, United States
Started: 2016

In graphics, you move data to the GPU, calculate for a long time, and send the result. Inference is the opposite: you move a massive volume of weights from memory to compute to calculate a single word, and then repeat. Traditional HBM memory bandwidth is too slow for this.

0
From the Full Interview

This answer is part of a full interview with Andrew Feldman, Co-founder & CEO at Cerebras Systems.

Share this Answer

Found this insight valuable? Share it with your network to help others learn from Andrew Feldman's experience.

Cite This Answer

Use this answer in your research, article, or academic work