Why do standard GPUs struggle with data movement during the decode phase of inference?
Replied byAndrew Feldman
Co-founder & CEO at Cerebras Systems
Niche: Technology
Revenue: Approx. $72 Million/month
Location: Sunnyvale, California, United States
Started: 2016
In graphics, you move data to the GPU, calculate for a long time, and send the result. Inference is the opposite: you move a massive volume of weights from memory to compute to calculate a single word, and then repeat. Traditional HBM memory bandwidth is too slow for this.
0
From the Full Interview
This answer is part of a full interview with Andrew Feldman, Co-founder & CEO at Cerebras Systems.
Share this Answer
Found this insight valuable? Share it with your network to help others learn from Andrew Feldman's experience.
Cite This Answer
Use this answer in your research, article, or academic work
Related Answers
What was your worst historical investment experience, and what risk management lesson did it impart?
By Sir Paul Marshall
Finance
Not Publicly Disclosed/mo
How would you describe your relationship with President Trump, and what are your thoughts on his policies?
By Bill Ackman
Finance
Not Publicly Disclosed/mo
What was your initial business model when you first launched Corgi?
By Emily Yuan
Finance
Approx. $3.3 Million/mo
What role did your early junior mining investments play in your career?
By Aaron Hoddinott
Finance
Not Publicly Disclosed /mo
What is your overarching outlook on macroeconomic cycles and the artificial intelligence revolution?
By Sir Paul Marshall
Finance
Not Publicly Disclosed/mo
What inspired you to write down the ten most influential people in your life?
By Nigel Morris
Finance
Not Publicly Disclosed/mo
How did you protect your mental health to avoid burning out during your career?
By Nigel Morris
Finance
Not Publicly Disclosed/mo