Why do standard GPUs struggle with data movement during the decode phase of inference?
Replied byAndrew Feldman
Co-founder & CEO at Cerebras Systems
Niche: Technology
Revenue: Approx. $72 Million/month
Location: Sunnyvale, California, United States
Started: 2016
In graphics, you move data to the GPU, calculate for a long time, and send the result. Inference is the opposite: you move a massive volume of weights from memory to compute to calculate a single word, and then repeat. Traditional HBM memory bandwidth is too slow for this.
0
From the Full Interview
This answer is part of a full interview with Andrew Feldman, Co-founder & CEO at Cerebras Systems.
Share this Answer
Found this insight valuable? Share it with your network to help others learn from Andrew Feldman's experience.
Cite This Answer
Use this answer in your research, article, or academic work
Related Answers
How does the 750-megawatt data center deal with OpenAI help scale their cloud footprint?
By Andrew Feldman
Technology
Approx. $72 Million/mo
Why do you find yourself arguing with large language models during the drafting process?
By Chris Best
Technology
Not Publicly Disclosed /mo
Why do so many businesses struggle with short-term thinking?
By Luke Sarsfield
Technology
/mo
How does the Netflix evolution analogy apply to the impact of fast inference?
By Andrew Feldman
Technology
Approx. $72 Million/mo
Why does Cerebras's wafer-scale engine process AI workloads faster than standard GPUs?
By Andrew Feldman
Technology
Approx. $72 Million/mo
What was the hardest packaging problem you had to solve during early chip building?
By Andrew Feldman
Technology
Approx. $72 Million/mo
What was the initial reaction and atmosphere in the mission control room during liftoff?
By Pawan Kumar Chandana
Technology
Pre-Revenue/mo