01 / Batch

Batch processing.

Asynchronous, always available. Ideal for document pipelines, data processing, and enterprise workloads.

Service Description Unit Price per Unit SLA
Document Summarization Extractive & abstractive summarization via flagship models Per page (~500 words) €0.012 4h turnaround
Text Classification Multi-label, sentiment, topic categorization Per item €0.003 1h turnaround
Entity Extraction NER, PII detection, structured data extraction Per 1K characters €0.005 1h turnaround
Embedding Generation BGE-M3 / Stella / GTE, 1024-3584d vectors Per 1K tokens €0.00008 30min turnaround
Batch RAG Pipeline Ingest → chunk → embed → index → retrieve → generate Custom quote Contact Per SLA
Batch Chat Completion Submit up to 50K prompts, results within SLA window Per 1M tokens See model pricing 8h turnaround
02 / Real-time

Real-time chat.

Premium capacity for interactive applications. Limited availability during startup phase.

Trial
For evaluation and prototyping
Free
  • 10 concurrent requests
  • Efficiency-tier models only
  • Rate-limited (5 req/min)
  • 10K tokens/month
  • Community support
  • No SLA
Get Started
Enterprise
For high-volume production
€0.09/min
  • 500 concurrent requests
  • Priority GPU scheduling
  • 99.95% uptime SLA
  • Dedicated inference pool
  • 24/7 priority support
  • Custom model fine-tuning
Contact Sales
Startup phase note: Real-time chat capacity is intentionally limited during our initial phase. As a Huawei Ascend partner, we prioritize batch processing workloads where throughput and data sovereignty matter more than latency. Real-time capacity will expand with Phase 2 infrastructure (Q4 2026).
03 / SLA

Service level agreement.

Uptime commitments, support response times, and incident handling.

Batch Processing SLA

MetricCommitment
Job acceptance99.9% — always accepting
Summarization turnaround≤ 4 hours
Classification turnaround≤ 1 hour
Embedding turnaround≤ 30 minutes
Batch chat turnaround≤ 8 hours
Missed SLA credit10% of job cost

Real-Time SLA

MetricBusinessEnterprise
Uptime99.5%99.95%
Max downtime/mo3.6 hours22 minutes
p95 latency< 3s< 2s
Support response< 4h email< 30min
Incident notification< 1 hour< 15 minutes
Uptime credit5% per 30min10% per 30min
Exclusions: SLA does not apply during scheduled maintenance windows (notified 72 hours in advance, conducted 02:00-05:00 CET). Force majeure events, customer-side network issues, and upstream model availability beyond our control are excluded. Full SLA document available upon contract signing.
99.95%
Enterprise Uptime Target
99.9%
Batch Availability
<2s
Enterprise p95 Latency
72h
Maint. Notice Period
04 / Calculator

Cost calculator.

Estimate your monthly spend based on model choice and token volume.

100K 50M 500M
Estimated Monthly Cost
€0.00/month
05 / Comparison

Market comparison.

How Nypples compares to alternative providers for self-hosted EU workloads.

Provider Data Residency DeepSeek V4 Pro Input/M Self-Hosted GDPR DPA No US Sub-processors
Nypples Industries EU (Trento) €1.60 Yes Yes Yes
DeepSeek API China ~$0.27 (~€0.25) No No Yes
Together AI US $1.74 (~€1.60) No Varies No
Fireworks AI US $1.50 (~€1.38) No Varies No
HuggingFace Inference US/EU $1.50 (~€1.38) Possible Enterprise Only No
Our pricing philosophy: We price competitively with the market — not a race to the bottom. What you pay for is data sovereignty. When you choose Nypples, your prompts and completions never leave the EU. No GDPR Schrems II concerns. No US CLOUD Act exposure. For regulated industries (healthcare, legal, insurance, government), this difference is existential.