01 / Catalog

Model catalog.

All models are self-hosted on Huawei Ascend infrastructure in Trento. Pricing in EUR, billed per 1M tokens.

Model Params Context Precision Real-Time Input / 1M tok Output / 1M tok
DeepSeek V4 Pro 1.6T / 49B active 1M FP4+FP8 Premium €1.60 €3.20
DeepSeek V4 Flash 284B / 13B active 1M FP4+FP8 Available €0.25 €0.55
GLM-5.2 744B / 40B active 256K FP4 Premium €1.30 €4.05
Kimi K2.7 Code 1.1T / ~40B active 256K FP8 Premium €0.90 €3.70
Kimi K3 Q4 2026 ~1.4T / ~46B active 512K FP4+FP8 Coming €1.10 €4.20
Qwen3.7-Max Q4 2026 ~390B / ~30B active 128K FP8 Coming €1.15 €3.50
Model Params Context Precision Real-Time Input / 1M tok Output / 1M tok
Qwen3.7-Plus ~70B 128K FP8 Available €0.30 €1.20
NVIDIA Nemotron 3 Ultra ~50B 128K FP8 Available €0.55 €3.30
Gemma 4 31B 31B 128K FP8 Available €0.35 €0.90
GPT-OSS 120B 120B 128K FP8 Batch Only €0.15 €0.55
Qwen3.5 9B 9B 128K FP8 Available €0.15 €0.23
Model Quantization VRAM/Node Min Nodes Throughput Gain Input / 1M tok Output / 1M tok
DeepSeek V4 Flash (INT4) GPTQ 4-bit ~85 GB 2 +35% €0.18 €0.40
GLM-5.2 (INT4) AWQ 4-bit ~200 GB 4 +30% €0.90 €2.80
Qwen3.7-Plus (Q4) GPTQ 4-bit ~35 GB 1 +50% €0.20 €0.85
Nemotron 3 Ultra (Q4) AWQ 4-bit ~28 GB 1 +50% €0.38 €2.30
About quantization: Quantized models use 4-bit weight compression (GPTQ/AWQ) to reduce memory footprint and increase throughput. Quality degradation is typically <2% on standard benchmarks while enabling 30-50% higher throughput per GPU node. Ideal for high-volume batch processing and cost-sensitive deployments.
Model Dimensions Max Tokens Batch Ready Price / 1M tokens
BGE-M3 1024d 8192 Yes €0.08
Stella-400M 1024d 8192 Yes €0.06
GTE-Qwen2-7B 3584d 32768 Yes €0.12
Jina-Embeddings-v3 1024d 8192 Yes €0.07
02 / Availability

Real-time availability tiers.

Concurrency and latency commitments per model tier.

Premium

Customer-facing chat

Limited real-time capacity. Suitable for customer-facing chat and interactive applications. Guaranteed latency under 2 seconds. Available for DeepSeek V4 Pro, GLM-5.2, and Kimi K2.7 Code.

Available

Internal tools & prototyping

Real-time access with moderate concurrency limits. Suitable for internal tools, prototyping, and moderate-traffic applications. Covers Flash variants and efficiency-tier models.

Batch only

No real-time SLA — unlimited volume

Jobs are queued and processed with priority scheduling. Ideal for document summarization, classification, embedding extraction, and large-scale RAG pipelines. Always available, no concurrency limits.