01 / Batch
Batch processing.
Asynchronous, always available. Ideal for document pipelines, data processing, and enterprise workloads.
| Service | Description | Unit | Price per Unit | SLA |
|---|---|---|---|---|
| Document Summarization | Extractive & abstractive summarization via flagship models | Per page (~500 words) | €0.012 | 4h turnaround |
| Text Classification | Multi-label, sentiment, topic categorization | Per item | €0.003 | 1h turnaround |
| Entity Extraction | NER, PII detection, structured data extraction | Per 1K characters | €0.005 | 1h turnaround |
| Embedding Generation | BGE-M3 / Stella / GTE, 1024-3584d vectors | Per 1K tokens | €0.00008 | 30min turnaround |
| Batch RAG Pipeline | Ingest → chunk → embed → index → retrieve → generate | Custom quote | Contact | Per SLA |
| Batch Chat Completion | Submit up to 50K prompts, results within SLA window | Per 1M tokens | See model pricing | 8h turnaround |
02 / Real-time
Real-time chat.
Premium capacity for interactive applications. Limited availability during startup phase.
Trial
For evaluation and prototyping
Free
- 10 concurrent requests
- Efficiency-tier models only
- Rate-limited (5 req/min)
- 10K tokens/month
- Community support
- No SLA
Business
For startups and SMBs
€0.06/min
- 100 concurrent requests
- All models incl. Flagship
- 99.5% uptime SLA
- Unlimited tokens
- Email support (4h)
- GDPR DPA included
Enterprise
For high-volume production
€0.09/min
- 500 concurrent requests
- Priority GPU scheduling
- 99.95% uptime SLA
- Dedicated inference pool
- 24/7 priority support
- Custom model fine-tuning
Startup phase note: Real-time chat capacity is intentionally limited during our initial phase. As a Huawei Ascend partner, we prioritize batch processing workloads where throughput and data sovereignty matter more than latency. Real-time capacity will expand with Phase 2 infrastructure (Q4 2026).
03 / SLA
Service level agreement.
Uptime commitments, support response times, and incident handling.
Batch Processing SLA
| Metric | Commitment |
|---|---|
| Job acceptance | 99.9% — always accepting |
| Summarization turnaround | ≤ 4 hours |
| Classification turnaround | ≤ 1 hour |
| Embedding turnaround | ≤ 30 minutes |
| Batch chat turnaround | ≤ 8 hours |
| Missed SLA credit | 10% of job cost |
Real-Time SLA
| Metric | Business | Enterprise |
|---|---|---|
| Uptime | 99.5% | 99.95% |
| Max downtime/mo | 3.6 hours | 22 minutes |
| p95 latency | < 3s | < 2s |
| Support response | < 4h email | < 30min |
| Incident notification | < 1 hour | < 15 minutes |
| Uptime credit | 5% per 30min | 10% per 30min |
Exclusions: SLA does not apply during scheduled maintenance windows (notified 72 hours in advance, conducted 02:00-05:00 CET). Force majeure events, customer-side network issues, and upstream model availability beyond our control are excluded. Full SLA document available upon contract signing.
99.95%
Enterprise Uptime Target
99.9%
Batch Availability
<2s
Enterprise p95 Latency
72h
Maint. Notice Period
04 / Calculator
Cost calculator.
Estimate your monthly spend based on model choice and token volume.
100K
50M
500M
Estimated Monthly Cost
€0.00/month
05 / Comparison
Market comparison.
How Nypples compares to alternative providers for self-hosted EU workloads.
| Provider | Data Residency | DeepSeek V4 Pro Input/M | Self-Hosted | GDPR DPA | No US Sub-processors |
|---|---|---|---|---|---|
| Nypples Industries | EU (Trento) | €1.60 | Yes | Yes | Yes |
| DeepSeek API | China | ~$0.27 (~€0.25) | No | No | Yes |
| Together AI | US | $1.74 (~€1.60) | No | Varies | No |
| Fireworks AI | US | $1.50 (~€1.38) | No | Varies | No |
| HuggingFace Inference | US/EU | $1.50 (~€1.38) | Possible | Enterprise Only | No |
Our pricing philosophy: We price competitively with the market — not a race to the bottom. What you pay for is data sovereignty. When you choose Nypples, your prompts and completions never leave the EU. No GDPR Schrems II concerns. No US CLOUD Act exposure. For regulated industries (healthcare, legal, insurance, government), this difference is existential.