Strata v0.1.42 Speeds Up 125B Model Runs on Consumer GPUs
Strata v0.1.42 adding roughly 13–16% faster decode for running the 125-billion-parameter Qwen3.8-Flash-Next model on ordinary hardware, with gains reaching up to 30% on cards whose CPU outruns their PCIe link. Two new defaults drive the speed: a "tail skip" that drops rarely-used experts and a smarter PCIe/CPU split measured from a live CPU reading. Those changes do shift answers slightly, though Strata measured the difference as tiny and offers a three-flag combo to restore exact 0.1.41 output. The release also folds in five long-delayed fixes spanning multi-GPU splits, Windows AMD, Intel Arc and AVX2-less CPUs, plus a grab bag of opt-in speedups, while leaving a few known issues unresolved.
Strata v0.1.42 Speeds Up 125B Model Runs on Consumer GPUs @ Linux Compatible
Strata v0.1.42 Speeds Up 125B Model Runs on Consumer GPUs
Strata v0.1.42 has improved the speed of running the 125-billion-parameter Qwen3.8-Flash-Next model on consumer GPUs by 13-16%, with performance gains of up to 30% on certain hardware configurations. This update introduces two key features: a "tail skip" that omits rarely-used experts and a more efficient PCIe/CPU split based on real-time measurements, although these changes may slightly alter outputs. Alongside speed enhancements, the release also includes several long-awaited fixes for multi-GPU setups, Windows AMD, Intel Arc, and CPUs lacking AVX2 support, while some known issues remain unresolved. Users can easily transition to this new version without affecting their existing models or configurations
