Nowadays we do billions of guesses per second on modern laptops, not even particularly high-end ones; so when this person gets one token every four seconds for a 2.8 TRILLION parameter LLM, using AN OLD MACBOOK PRO, to me it suggests that local LLMs will soon be everywhere:
Leave a Reply