120B-Class AI

Running Local

128GB Unified Memory Gives N5 MAX The Headroom To Run 120B-Class Open Models Locally, While Ryzen AI Max+ 395 And Radeon 8060S Keep Reasoning, Coding, And Agent Workflows Moving At Responsive Speeds.

gpt-oss-120b

120B-Class Open Reasoning, Running Local

59.30 tokens/s
Text Generation
668.79 tokens/s
Prompt Processing
59.03 GB
MXFP4

Qwen3.5-122B-A10B

122B Multimodal Intelligence, On Device

31.76 tokens/s
Text Generation
317.58 tokens/s
Prompt Processing
69.11 GB
Q4_K_M

gpt-oss-20b

Fast Local Reasoning for Everyday Agents

83.09 tokens/s
Text Generation
1,575.75 tokens/s
Prompt Processing
11.28 GB
MXFP4

Gemma 4 26B-A4B

Efficient Intelligence, Responsive by Design

70.33 tokens/s
Text Generation
1,200.58 tokens/s
Prompt Processing
15.64 GB
Q4_K_M

*Benchmark Note
Internal test results based on GGUF quantized models using llama.cpp llama-bench, with 99 GPU layers and Flash Attention enabled. PP512 measures prompt-processing throughput; TG128 measures text-generation speed. Actual performance may vary by model version, quantization, context length, system configuration, and software environment.