Running a 35B LLM at 128K Context on Just €870 of Used Hardware

A recent Medium article details how a developer successfully runs a 35 billion parameter large language model with an impressive 128K context window entirely on local hardware. The total cost of the used equipment amounts to just €870, proving that high-performance AI inference is becoming increasingly accessible without enterprise-level budgets.

The setup eliminates any dependence on cloud computing services, which typically impose ongoing costs and require consistent internet connectivity. By running everything locally, users gain full control over their data privacy and avoid the recurring expenses that come with cloud-based API access to similar-sized models.

This achievement signals a broader trend in the AI community where optimized software and affordable second-hand components are closing the gap between consumer-grade setups and professional infrastructure. As model quantization and inference techniques continue to improve, more developers and hobbyists can experiment with capable LLMs right from their own desks.

Read More at the original source →