Strata v0.1.40.4 Restores Decode Speed on Pascal GPUs After Regression

Published by

Strata v0.1.40.4 is a hotfix that restores decode performance for older NVIDIA Pascal GPUs, such as the GTX 1080 and Tesla P40, which experienced a significant slowdown due to a previous commit that incorrectly applied changes across all CUDA architectures. This regression was identified and addressed by community members who conducted thorough testing, leading to a two-line fix that reinstated the essential compiler hints and corrected a prefetch flag issue. Users with Pascal GPUs can now expect their token speeds to return to normal levels, while those with newer RTX models will see no changes, as the compiled code remains the same. Installation of the update is straightforward for Linux users, requiring adequate system resources to run the model effectively



Strata v0.1.40.4 Restores Decode Speed on Pascal GPUs After Regression

Strata v0.1.40.4 delivers a critical hotfix that restores decode performance for older NVIDIA Pascal GPUs, including the GTX 1080 and Tesla P40, after a regression halved token speeds. The issue stemmed from commit 2e4ddf6, which removed essential __restrict__ compiler hints and added a non-overlappable prefetch flag to every architecture rather than just high-end silicon. Maintainer Niko lacks Pascal hardware, so community members paulhothersall and lineape conducted rigorous bisects and A/B testing to pinpoint the cause and validate the two-line solution. The patch targets cards below sm_70 exclusively, meaning users with RTX 20, 30, 40, or 50 series GPUs will see no behavioral change as the compiled code remains byte-identical.

Strata v0.1.40.4 Restores Decode Speed on Pascal GPUs After Regression @ Linux Compatible