Speaking during an interview on CNBC’s “Squawk on the Street” segment earlier this week, CEO of cybersecurity giant Palo Alto Networks Nikesh Arora implored the tech industry to lower the cost of AI.
During the segment, the chief executive argued that the cost to use large language models (LLMs) has to drop by 20 percent by 2027 — and 90 percent by 2028 — for the tech to be useful to enterprises.
“We need to see the pricing for AI come down,” Arora said.



30B models are no featherweight at all and an old 3080 won’t be able to lift them due to insufficient RAM and even if it did, low bandwidth would make it so you’d only be able generate a few dozens of tokens per minute.
This message alone would take like 2 minutes to generate.
A quantized 30b model will run on a 3080. I know because I’ve done it. But yeah, it needs to be quantized really small.
Edit: wait, no, you’re right, I’m thinking of a 3090. I’ve done it on a 3090, not a 3080. Quantized to 4 bits. And yes, it’s very slow. That’s enough to replace a human, right?