AI Compute Server | 128 GB GDDR7 | EPYC 512GB

687 519,83 Kč
687 CZK excluding VAT

Ideal for AI studios , advertising agencies, and generative AI developers who need to generate photorealistic videos and images locally — without cloud costs, no API limits, and no data leaks. One server replaces dozens of Mac minis or other AI workstations at a fraction of the total cost

Last piece in stock

Best price guarantee!
Add to Cart
AI Compute Server | 128 GB GDDR7 | EPYC 512GB
AI Compute Server | 128 GB GDDR7 | EPYC 512GB
687 519,83 Kč
687 CZK excluding VAT

Price without VAT when using our hosting
🎓 Free training for beginners and advanced
🛡️ 24 month warranty on all devices
Express service – quick repair in our center

Secure card payment

 

Price with VAT (21%)
CZK831 899
Without VAT: 687 520 CZK  · invoice with VAT number / VAT number · delivery within 7 days
4× RTX 5090 Blackwell — 128GB GDDR7 VRAM total
Burn-in 12 minutes, all GPUs 575W — 0 errors
AI stack pre-installed (vLLM, llama.cpp, PyTorch, CUDA)
BMC remote management IPMI 2.0 — no physical access
Ready to go — just connect and log in
✓ Burn-in PASS Blackwell sm_120 CUDA 12.8 Order #KE21013 AMD EPYC 7643 512 GB ECC RAM
128 GB
Total GPU VRAM
~ 880 T
FP16 TFLOPS
1 631 t/s
LLM batch
0
Burn-in errors
AI content generation

Server for generating video and images indistinguishable from reality

128GB of GDDR7 VRAM across four Blackwell GPUs lets you run the largest open-source diffusion and video models at full accuracy — without offloading, quantization, or waiting. No cloud, no API fees, no data leaving your network.

Pre-installed and tested
🎬

Video generation — HunyuanVideo, Wan2.1, CogVideoX

128 GB VRAM allows for the generation of long sequences in high resolution without tiling or frame-splitting limitations. Full CUDA 12.8 support for all modern video models.

128 GB VRAM
🖼️

Image generation

FLUX.1, SD3.5, Stable Cascade — full precision FP16, batch, ControlNet, IP-Adapter.

🧠

LLM Inference for Teams

vLLM tensor-parallel × 4 GPU — 1,631 threads/s for 64 concurrent users. Qwen, Llama, Mistral, DeepSeek.

Hardware

Complete technical specifications

GPU (× 4 pieces)
Door DesignNVIDIA GeForce RTX 5090
ArchitectureBlackwell, sm_120
VRAM on GPU32 GB GDDR7
Total VRAM128 GB
Multi-process streaming.170 SM/card
Memory bandwidth~1,530 GB/s / GPU
TDP575W/GPU (4x = 2300W)
PCIeGen4 × 16 per slot
CPU and system memory
CPUAMD EPYC 7643
Cores/Fibers48C / 96T
Max. frequency3,64 GHz
L3 cache256 MiB
System RAM512 GB DDR4-2666
ECCYes (RDIMM)
DIMMs8× 64GB Samsung
Memory channels8 (fully equipped)
Motherboard and storage
MotherboardASRockRack ROMED8-2T/BCM
BIOSP4.30 (20. 5. 2026)
NVMe SSDsCrucial P310 2TB
Sequential reading7 MB / s
Sequential writing6 MB / s
Random 4K IOPS467 K / 477 K black/white
Sew2× 10GbE + IPMI/BMC
Remote controlASRock Rack BMC 3.08
Software (pre-installed)
OSUbuntu LTS 24.04.4
NVIDIA driver595.71.05 (open)
CUDA driver/toolkit13.2 / 12.8
cuDNN9.10.02
PyTorch2.9.1+cu128
vLLM0.14.0
call.cppCUDA build, sm_120
Python venv/home/logic/llm-env
Measured power

Real benchmarks from this particular machine

Measured on 10/7/2026 on host dm2, order #KE21013. Numbers are measured, not given by the manufacturer.

FP16 performance on GPU (8192² matmul)

GPU 0
219,1 T
GPU 1
221,3 T
GPU 2
219,4 T
GPU 3
218,3 T

Total ~880 TFLOPS FP16 · +33% per card vs. RTX 4090

VRAM transfer rate (GDDR7)

GPU 0
1,533 GB/s
GPU 1
1,533 GB/s
GPU 2
1,533 GB/s
GPU 3
1,533 GB/s

PCIe host↔GPU: 27,6 / 28,3 GB/s (H2D / D2H)

vLLM · Qwen2.5-32B-Instruct AWQ · tensor-parallel × 4 · 64 concurrent users

Batch aggregated throughput64 concurrent users
1 631 flow/s
Door DesignQwen2.5-32B AWQ (marlin kernel)
32B
GPU connectionTensor-parallel × 4 GPUs
✓ all 4

llama.cpp · Qwen2.5-14B Q4_K_M · 1 GPU · llama-bench

Generating tokenstg128 benchmark
147 flow/s
Prompt processingpp512 benchmark
8 234 flow/s
QuantizationQ4_K_M, 8,37 GiB
14,8B
7 074 MB / s
Reading sec.
6 341 MB / s
Sec. entry
467 K
4K Read IOPS
477 K
4K Write IOPS
Reliability

12-minute burn-in at full load — 0 errors

GPU burn — 720 s sustained FP16 matmul (8192×8192), all 4 GPUs at 575 W ✓ PASS
GPU 0
84°C
575W 2,4GHz
0 errors
GPU 1
88°C
575W 2,35GHz
0 errors
GPU 2
83°C
575W 2,5GHz
0 errors
GPU 3
80°C
575W 2,5GHz
0 errors
Throttling threshold RTX 5090 ≈ 90 °C · maximum measured 88 °C (GPU 1) · no performance throttling · SM clock 2,3–2,5 GHz at all times · combined GPU + 96 CPU threads test also passed without errors · kernel log: no Xid / NVRM / PCIe AER errors
Comparison with the competition

Why our server instead of the Mac Studio M3 Ultra 512 GB?

The Mac Studio M3 Ultra in the 512 GB / 8 TB configuration is the most powerful single-box Apple machine for AI. However, it costs ~1,000,000 CZK — for this price you can buy more than one of our servers.

Mac Studio M3 Ultra 512 GB / 8 TB Apple BTO · 32-core CPU · 80-core GPU · macOS only · no CUDA · ~819 GB/s bandwidth
Mac Studio M3 Ultra
~1,000,000 CZK
BTO configuration 512GB / 8TB
Our server (including VAT)
831 899 CZK
4x more CUDA performance, lower price
ParameterMac Studio M3 Ultra 512GB4× RTX 5090 Server (KENTINO)
Price with VAT~1,000,000 CZK more expensive831 899 CZK winner
GPU accelerator80-core M3 Ultra GPUs · ~40–60 TFLOPS FP16 · Metal only4× RTX 5090 Blackwell · ~880 TFLOPS FP16 × 15+ · CUDA 12.8
VRAM / AI memory512 GB unified (CPU+GPU shared) 819 GB/s128 GB GDDR7 · ~1530 GB/s per GPU 5x faster + 512GB ECC RAM
LLM speed 32B~35–40 flow/s · 1 user1,631 flow/s batch × 40+ · 64 users
Video generationMetal unstable for CUDA modelsCUDA · HunyuanVideo · Wan2.1 · native support
Concurrent users1–3 practically64+ (in LLM tensor parallel)
AI ecosystemmacOS / Metal / MLX · without CUDALinux / CUDA · PyTorch · fine-tuning · all
Remote controlSSH · no IPMI · physical access for rebootIPMI 2.0 · BMC · iKVM console · remote BIOS
Fine-tuningMPS unstable · not recommendedFull CUDA LoRA QLoRA FSDP

For the price of a Mac Studio M3 Ultra, you get an RTX 5090 server with the CUDA ecosystem

The Mac Studio 512 GB is great for single-user workloads with huge models in unified memory. But it's more expensive, runs on macOS without CUDA, can handle 1-3 users max, and doesn't support fine-tuning. Our server is faster for batch inference, has full CUDA compatibility, and can handle the whole team at once.

~15x more TFLOPS FP16 64+ concurrent users Full CUDA ecosystem Fine-tuning, LoRA, training BMC remote management 24/7
Pre-installed software

No installation — just plug and log in

NVIDIA Driver + CUDA
595.71.05 · CUDA 13.2 / toolkit 12.8 · cuDNN 9.10
PyTorch 2.9.1
+cu128 · sm_120 Blackwell · NCCL multi-GPU
vLLM 0.14.0
tensor-parallel · AWQ marlin · OpenAI API compatible
llama.cpp (CUDA build)
sm_120 · commit c4ae9a8
HuggingFace stack
transformers · accelerate · diffusers · hub
Python venv
/home/logic/llm-env · models in /home/logic/models

⚡ Delivery and first launch

  • Delivery within 7 working days after order confirmation
  • Invoice with VAT number — VAT deductible for companies
  • Change the system password and BMC password immediately
  • Set the BMC to a static IP on your network
  • Configure the firewall before running API inference
  • SSH user: logic (sudo) · BMC user: admin
  • GPU power consumption can be regulated: nvidia-smi -pl [watts]
  • Technical support KENTINO sro — e-mail + phone
Technical notes All GPUs run on PCIe Gen4 × 16 — AMD EPYC 7003 / ROMED8-2T platform is Gen4 (no impact on LLM inference performance). RTX 5090 does not have NVLink — multi-GPU communication over PCIe P2P ~24,3 GB/s (standard for this class). ECC GPU memory is not available on consumer RTX 5090 (expected, ECC is on system RAM). Benchmarks measured 7/10/2026, part #KE21013, dm2 host. All numbers are measured, not reported by the manufacturer.

Producer

Pcpraha

Gaming computers of the Pc Praha brand. The best solution in the Czech Republic and the EU. Pro gaming computers and mining rigs and Asici.

Questions and answers Q & A

Ask a question
There are no questions yet