01 / Compute
Europe's first Ascend GPU cluster.
Purpose-built for AI inference at scale. No US cloud, no NVIDIA tax, no data export.
20
Ascend 910C Nodes
320
TFLOPS FP16 per Node
6.4
PFLOPS Total Compute
800
Gbps Interconnect
Huawei Ascend 910C: Each node features the Da Vinci architecture with native FP16 performance of approximately 320 TFLOPS. Paired with 64 GB HBM2E memory per GPU, delivering 1.6 TB/s memory bandwidth. The cluster is organized into four 5-node groups connected via 800 Gbps RoCE v2 networking, providing near-InfiniBand latency for tensor parallelism across nodes.
02 / Specifications
Ascend 910C vs NVIDIA H100.
Side-by-side. We don't hide from the comparison.
| Specification | Ascend 910C | NVIDIA H100 (for reference) |
|---|---|---|
| FP16 Performance | ~320 TFLOPS | ~350 TFLOPS (less FP8) |
| INT8 Performance | ~640 TOPS | ~700 TOPS |
| Memory | 64 GB HBM2E | 80 GB HBM3 |
| Memory Bandwidth | 1.6 TB/s | 3.35 TB/s |
| Interconnect | HCCS + 800G RoCE | NVLink + InfiniBand |
| Software Stack | CANN 8.0 / MindSpore | CUDA 12 / cuDNN |
| Partner Price per Node | ~€29,000 | ~€37,000+ |
03 / Software Stack
Ascend ecosystem bridged to open-source AI.
Each layer is independently swappable and links to its upstream project.
06
Client SDKs · Web Console · API Keys
OpenAI-compatible REST & gRPC endpoints. Drop-in replacement for the OpenAI SDK.
OpenAI API
05
Kubernetes — Orchestration & Autoscaling
Pod-level GPU scheduling, HPA on queue depth, rolling model swaps without downtime.
K8s 1.30
04
vLLM Adapter — OpenAI-Compatible Serving
PagedAttention, continuous batching, and tensor parallelism across Ascend nodes.
vLLM
03
MindSpore Runtime
Huawei's native ML framework — graph-mode execution tuned for the Da Vinci architecture.
MindSpore
02
CANN 8.0 — Compute Architecture for Neural Networks
Huawei's CUDA-equivalent. Driver, runtime, and operator libraries for Ascend.
CANN 8.0
01
Ascend 910C Nodes — 20x Physical Hardware
Da Vinci architecture, 320 TFLOPS FP16 per node, 64 GB HBM2E, deployed in Trento, Italy.
Hardware
200 TB NVMe · Ceph
2 PB Object · MinIO
Prometheus · Grafana
Why this stack: The Ascend ecosystem replaces NVIDIA's CUDA + cuDNN with CANN + MindSpore. We bridge it to the open-source AI world via the vLLM Ascend adapter, exposing a fully OpenAI-compatible API. You write standard client code; we translate it onto sovereign European hardware.
04 / Budget
€1M angel round allocation.
Where every euro goes. Transparent by design.
Angel Round
€1.0M
20 nodes · 18+ mo runway
GPU Hardware — 20x Ascend 910C + servers
€580,00058%
Networking — 800G switches, optics, cabling
€85,0008.5%
Storage & Racks — NVMe + MinIO
€95,0009.5%
Colocation — Trento DC, Year 1
€120,00012%
Staff — 3 engineers, 6 months
€90,0009%
Contingency / Operations
€30,0003%
| Category | Amount | % of Budget | Details |
|---|---|---|---|
| GPU Hardware | €580,000 | 58% | 20x Ascend 910C modules + server chassis via Huawei EU partner program |
| Networking | €85,000 | 8.5% | 800G RoCE switches, optical transceivers, DAC cables |
| Storage & Racks | €95,000 | 9.5% | 200TB NVMe, 2PB MinIO object storage, rack enclosures |
| Colocation | €120,000 | 12% | Trento datacenter, 3 racks, 120kW power + cooling, Year 1 |
| Staff | €90,000 | 9% | 3 senior engineers (infra, ML ops, backend) for 6 months |
| Buffer | €30,000 | 3% | Contingency, legal, insurance, operational overhead |
05 / ROI
Return on investment.
Projected financials based on realistic utilization scenarios.
Burn rate
€29,000 / month
Colocation €10K, staff €15K, bandwidth & power €4K. Operational cost after Phase 1 deployment. Fully loaded.
Break-even
Month 8–10
At 25% GPU utilization across batch + real-time services, projected revenue reaches approximately €45,000/month.
Steady state
€95,000+ / month at 60% utilization
At target utilization, monthly revenue projected at €95,000+. Runway from angel round exceeds 18 months at current burn rate. Series A target: Q3 2027, €3–5M for multi-region EU expansion.