$16,000 Quad AI Geekom mini PC cluster gets DeepSeek V4 Flash treatment with 512GB RAM — reaches 14.61 tokens per second
GEEKOM squeezes a 250K-token AI workload into four small boxes
- Four GEEKOM A9 Mega mini PCs link together through simple USB4 cables
- Each unit runs on the AMD Ryzen AI Max+ 395 processor chip
- Combined cluster memory reaches 512GB across all four connected systems
Geekom has released the DeepSeek V4 Flash model across four A9 Mega mini PCs, forming a distributed cluster aimed at enterprise AI workloads.
The setup connects the devices through USB4 rather than relying on a traditional data center server.
Each A9 Mega runs on the AMD Ryzen AI Max+ 395 chip, which combines 16 Zen 5 CPU cores with Radeon 8060S graphics and unified memory in one compact chassis
Local processing over cloud dependence
The four-node configuration brings a combined 512GB of RAM to the task, with Ubuntu, ROCm, and DwarfStar software distributing the optimized model across the machines.
An OpenAI-compatible API links applications and AI agents to the cluster, while USB4 removes any need for a proprietary switch or server rack.
The arrangement allows organizations to keep prompts, documents, source code, and credentials within local infrastructure rather than routing them through a public cloud.
Businesses could theoretically build private knowledge assistants capable of searching contracts, manuals, and internal reports without external exposure.
Sign up to the TechRadar Pro newsletter to get all the top news, opinion, features and guidance your business needs to succeed!
For agent-based systems, Geekom says the cluster can process tools, policies, memory, and logs before executing any action, with a technical path that has operated using contexts up to 250K tokens.
Performance numbers under scrutiny
Testing reported by Geekom recorded approximately 14.61 tokens per second at single concurrency, while P95 time to first token reached about 0.42 seconds.
The company says the configuration also provides greater capacity for long prompts, rather than concentrating solely on faster short-response generation.
Users can begin with one or two A9 Mega systems before expanding the configuration to four nodes as workloads increase over time.
Each machine can operate independently, while connected systems can contribute to distributed inference when greater computing capacity becomes necessary for demanding workloads.
The A9 Mega mini PC itself supports up to 128GB of LPDDR5x memory, although four systems provide the stated 512GB cluster capacity for distributed inference workloads.
Its specifications include 120W sustained performance, dual M.2 PCIe Gen4x4 storage slots, and support for up to 8TB of RAID-configured storage.
Geekom lists the A9 Mega at $3,999 for a configuration with 128GB RAM and 2TB SSD, while the four-system cluster reaches $16,000.
The individual units also carry 126 total TOPS of combined processing output, with 120W of sustained thermal design power under load.
It also supports Wi-Fi 7, Bluetooth 5.4, dual 2.5G Ethernet, and four simultaneous 8K displays for professional environments.
The performance values above are vendor-supplied benchmarks without third-party testing - therefore, the real-world reliability of the setup for sustained enterprise use has not been independently confirmed.
Follow TechRadar on Google News and add us as a preferred source to get our expert news, reviews, and opinion in your feeds.

Efosa has been writing about technology for over 7 years, initially driven by curiosity but now fueled by a strong passion for the field. He holds both a Master's and a PhD in sciences, which provided him with a solid foundation in analytical thinking.
You must confirm your public display name before commenting
Please logout and then login again, you will then be prompted to enter your display name.