Strata v0.1.38 Runs a 125-Billion-Model LLM on Any Gaming PC

Published by

Strata v0.1.38 is an open-source inference engine that allows users to run a 125-billion-parameter language model on standard gaming PCs, utilizing system resources like GPU, CPU, RAM, and SSD. The latest update includes a critical security fix to prevent DNS rebinding and cross-site attacks by restricting server responses to trusted hosts. Performance testing shows that the engine can generate up to 120 tokens per second, with an average of 75 tokens per second, while operating entirely locally without any cloud dependency. Additionally, the update enhances compatibility and speed for both NVIDIA and AMD graphics cards, making it accessible for users with varying hardware capabilities



Strata v0.1.38 Runs a 125-Billion-Model LLM on Any Gaming PC

Strata v0.1.38 is an open-source inference engine that spreads the 125-billion-parameter Qwen3.8-Flash-Next model across your GPU, CPU, RAM, and SSD so it runs on an ordinary gaming PC with a 12 GB card. The release's biggest change is a security fix that blocks DNS-rebinding and cross-site attacks against its local server by rejecting requests from untrusted hosts. In hands-on testing on Debian with ROCm and a Radeon 7900 XTX, output reached up to 120 tokens per second with a realistic 75 per-second average. Everything runs locally and privately, with no cloud dependency and nothing charged per token after purchase.

Strata v0.1.38 Runs a 125-Billion-Model LLM on Any Gaming PC @ Linux Compatible