Running a 753B-Parameter Model on a Single Workstation GPU: FreeToken Breaks the Parameter Ceiling for Local Inference
FreeToken, a bandwidth-adaptive MoE inference system developed by UC Berkeley and UT Austin researchers, enables a single workstation GPU to run a 753-billion-parameter model (Z.ai's GLM-5.2). The open-sourced system delivers 1.5-2.3x throughput gains over existing local inference engines while keeping time-to-first-token within 44 seconds.