01 / Catalog
Model catalog.
All models are self-hosted on Huawei Ascend infrastructure in Trento. Pricing in EUR, billed per 1M tokens.
| Model | Params | Context | Precision | Real-Time | Input / 1M tok | Output / 1M tok |
|---|---|---|---|---|---|---|
| DeepSeek V4 Pro | 1.6T / 49B active | 1M | FP4+FP8 | Premium | €1.60 | €3.20 |
| DeepSeek V4 Flash | 284B / 13B active | 1M | FP4+FP8 | Available | €0.25 | €0.55 |
| GLM-5.2 | 744B / 40B active | 256K | FP4 | Premium | €1.30 | €4.05 |
| Kimi K2.7 Code | 1.1T / ~40B active | 256K | FP8 | Premium | €0.90 | €3.70 |
| Kimi K3 Q4 2026 | ~1.4T / ~46B active | 512K | FP4+FP8 | Coming | €1.10 | €4.20 |
| Qwen3.7-Max Q4 2026 | ~390B / ~30B active | 128K | FP8 | Coming | €1.15 | €3.50 |
| Model | Params | Context | Precision | Real-Time | Input / 1M tok | Output / 1M tok |
|---|---|---|---|---|---|---|
| Qwen3.7-Plus | ~70B | 128K | FP8 | Available | €0.30 | €1.20 |
| NVIDIA Nemotron 3 Ultra | ~50B | 128K | FP8 | Available | €0.55 | €3.30 |
| Gemma 4 31B | 31B | 128K | FP8 | Available | €0.35 | €0.90 |
| GPT-OSS 120B | 120B | 128K | FP8 | Batch Only | €0.15 | €0.55 |
| Qwen3.5 9B | 9B | 128K | FP8 | Available | €0.15 | €0.23 |
| Model | Quantization | VRAM/Node | Min Nodes | Throughput Gain | Input / 1M tok | Output / 1M tok |
|---|---|---|---|---|---|---|
| DeepSeek V4 Flash (INT4) | GPTQ 4-bit | ~85 GB | 2 | +35% | €0.18 | €0.40 |
| GLM-5.2 (INT4) | AWQ 4-bit | ~200 GB | 4 | +30% | €0.90 | €2.80 |
| Qwen3.7-Plus (Q4) | GPTQ 4-bit | ~35 GB | 1 | +50% | €0.20 | €0.85 |
| Nemotron 3 Ultra (Q4) | AWQ 4-bit | ~28 GB | 1 | +50% | €0.38 | €2.30 |
About quantization: Quantized models use 4-bit weight compression (GPTQ/AWQ) to reduce memory footprint and increase throughput. Quality degradation is typically <2% on standard benchmarks while enabling 30-50% higher throughput per GPU node. Ideal for high-volume batch processing and cost-sensitive deployments.
| Model | Dimensions | Max Tokens | Batch Ready | Price / 1M tokens |
|---|---|---|---|---|
| BGE-M3 | 1024d | 8192 | Yes | €0.08 |
| Stella-400M | 1024d | 8192 | Yes | €0.06 |
| GTE-Qwen2-7B | 3584d | 32768 | Yes | €0.12 |
| Jina-Embeddings-v3 | 1024d | 8192 | Yes | €0.07 |
02 / Availability
Real-time availability tiers.
Concurrency and latency commitments per model tier.
Premium
Customer-facing chat
Limited real-time capacity. Suitable for customer-facing chat and interactive applications. Guaranteed latency under 2 seconds. Available for DeepSeek V4 Pro, GLM-5.2, and Kimi K2.7 Code.
Available
Internal tools & prototyping
Real-time access with moderate concurrency limits. Suitable for internal tools, prototyping, and moderate-traffic applications. Covers Flash variants and efficiency-tier models.
Batch only
No real-time SLA — unlimited volume
Jobs are queued and processed with priority scheduling. Ideal for document summarization, classification, embedding extraction, and large-scale RAG pipelines. Always available, no concurrency limits.