Server for generating video and images indistinguishable from reality
128GB of GDDR7 VRAM across four Blackwell GPUs lets you run the largest open-source diffusion and video models at full accuracy — without offloading, quantization, or waiting. No cloud, no API fees, no data leaving your network.
Video generation — HunyuanVideo, Wan2.1, CogVideoX
128 GB VRAM allows for the generation of long sequences in high resolution without tiling or frame-splitting limitations. Full CUDA 12.8 support for all modern video models.
Image generation
FLUX.1, SD3.5, Stable Cascade — full precision FP16, batch, ControlNet, IP-Adapter.
LLM Inference for Teams
vLLM tensor-parallel × 4 GPU — 1,631 threads/s for 64 concurrent users. Qwen, Llama, Mistral, DeepSeek.
Complete technical specifications
Real benchmarks from this particular machine
Measured on 10/7/2026 on host dm2, order #KE21013. Numbers are measured, not given by the manufacturer.
FP16 performance on GPU (8192² matmul)
Total ~880 TFLOPS FP16 · +33% per card vs. RTX 4090
VRAM transfer rate (GDDR7)
PCIe host↔GPU: 27,6 / 28,3 GB/s (H2D / D2H)
vLLM · Qwen2.5-32B-Instruct AWQ · tensor-parallel × 4 · 64 concurrent users
llama.cpp · Qwen2.5-14B Q4_K_M · 1 GPU · llama-bench
12-minute burn-in at full load — 0 errors
Why our server instead of the Mac Studio M3 Ultra 512 GB?
The Mac Studio M3 Ultra in the 512 GB / 8 TB configuration is the most powerful single-box Apple machine for AI. However, it costs ~1,000,000 CZK — for this price you can buy more than one of our servers.
| Parameter | Mac Studio M3 Ultra 512GB | 4× RTX 5090 Server (KENTINO) |
|---|---|---|
| Price with VAT | ~1,000,000 CZK more expensive | 831 899 CZK winner |
| GPU accelerator | 80-core M3 Ultra GPUs · ~40–60 TFLOPS FP16 · Metal only | 4× RTX 5090 Blackwell · ~880 TFLOPS FP16 × 15+ · CUDA 12.8 |
| VRAM / AI memory | 512 GB unified (CPU+GPU shared) 819 GB/s | 128 GB GDDR7 · ~1530 GB/s per GPU 5x faster + 512GB ECC RAM |
| LLM speed 32B | ~35–40 flow/s · 1 user | 1,631 flow/s batch × 40+ · 64 users |
| Video generation | Metal unstable for CUDA models | CUDA · HunyuanVideo · Wan2.1 · native support |
| Concurrent users | 1–3 practically | 64+ (in LLM tensor parallel) |
| AI ecosystem | macOS / Metal / MLX · without CUDA | Linux / CUDA · PyTorch · fine-tuning · all |
| Remote control | SSH · no IPMI · physical access for reboot | IPMI 2.0 · BMC · iKVM console · remote BIOS |
| Fine-tuning | MPS unstable · not recommended | Full CUDA LoRA QLoRA FSDP |
For the price of a Mac Studio M3 Ultra, you get an RTX 5090 server with the CUDA ecosystem
The Mac Studio 512 GB is great for single-user workloads with huge models in unified memory. But it's more expensive, runs on macOS without CUDA, can handle 1-3 users max, and doesn't support fine-tuning. Our server is faster for batch inference, has full CUDA compatibility, and can handle the whole team at once.
No installation — just plug and log in
⚡ Delivery and first launch
- Delivery within 7 working days after order confirmation
- Invoice with VAT number — VAT deductible for companies
- Change the system password and BMC password immediately
- Set the BMC to a static IP on your network
- Configure the firewall before running API inference
- SSH user: logic (sudo) · BMC user: admin
- GPU power consumption can be regulated: nvidia-smi -pl [watts]
- Technical support KENTINO sro — e-mail + phone









