Calibrated plans for deterministic AI.
From exploratory model stress testing to production failure clustering and fine-tuning remediation datasets. Zero hidden surcharges.
Trial
Experience Tanvelo's diagnostic architecture and semantic failure clustering on your first model.
*Trial ends on first limit hit. Upgrade to continue.
Starter
For active engineering teams running frequent model stress tests, comparisons, and root-cause remediation.
Pay with Razorpay • Cards, UPI & International
Pro
For high-velocity AI teams demanding unlimited projects, multi-seat collaboration, and audit-ready reports.
Pay with Razorpay • Cards, UPI & International
Pay securely in USD ($) or INR (₹) using Visa, Mastercard, American Express, Apple Pay, Discover, Diners Club, or domestic UPI.
*Trial ends on first limit hit. Upgrade to continue.
When you hit 50 tests or reach 5 days, benchmark execution pauses. All historical logs, reports, and clusters are preserved indefinitely.
Side-by-side tier specification
Compare diagnostic limits, evaluation depth, export rights, and queue priorities across every plan.
| Features & LimitsEvaluation Capabilities | Trial $0 5 days | Starter $20 Most Popular | Pro $50 5 Seats |
|---|---|---|---|
tuneLimits & Capacity | |||
Tests Total automated diagnostic test evaluations executed per billing period. | 50 tests / 5 days | 5,000 / month | 10,000 / month |
Projects Independent application or workspace containers to isolate models and datasets. | 1 | 3 | Unlimited |
Models Connected LLM endpoints (OpenAI, Ollama, vLLM, OpenRouter, or custom APIs) per project. | 1 | 2 per project | 5 per project |
Seats Team members with synchronized workspace access and shared diagnostic reports. | 1 | 1 | 5 |
Max per run Maximum parallel benchmark test cases executed in a single evaluation run. | 50 | 300 | 500 |
psychologyBenchmarking & Evaluation | |||
Categories Diagnostic failure domains: Reasoning, Grounding, Coding, Safety, plus Drift, Jailbreak, Edge Cases, etc. | 4 core | All 10 + custom requirements | All + bulk + custom criteria |
Evaluation Decoupled judging architecture with calibrated rubrics and automated safety checks. | LLM judge + Safety | Full evaluation | Full evaluation |
Diagnosis Semantic clustering of model failures into actionable root causes. | Score + top 3 | Full root-cause + recommendations | Full + prioritized |
datasetDatasets, Exports & Retention | |||
Datasets Supervised fine-tuning pairs (input prompt + verified gold completion) ready to download. | Preview | 500 / mo | 2,000 / mo |
Exports Available export formats and model regression comparison capabilities. | Dashboard | JSON + CSV + compare 2 models | + PDF + audit |
History Duration for which full benchmark logs, failure clusters, and generated traces are retained. | 5 days | 90 days | 12 months |
speedExecution Speed & Queue | |||
Speed Runner queue dispatch priority and concurrent worker allocation. | Standard queue | Priority queue | Fastest queue, small runs first |
| Action | Start free | ||
Use the toggle buttons above to switch plans
Need custom VPC runners, private judge rubrics, or 100K+ monthly tests?
Tanvelo offers dedicated isolated test runners inside your cloud perimeter (AWS, GCP, Azure), custom domain-specific synthetic generation pipelines, custom evaluation judges, and SLAs for mission-critical deployments.
Frequently Asked Questions
Clear terms and explanations regarding quotas, limits, international cards, and evaluations.
What does 'Trial ends on first limit hit. Upgrade to continue' mean?keyboard_arrow_down
The Trial is completely free ($0) with no credit card required. You receive 50 tests or 5 days of access (whichever comes first) on 1 connected model and 1 project. Once you exhaust 50 tests or the 5 days expire, benchmark execution pauses. All your historical logs, score cards, and clusters remain safely preserved, and you can instantly upgrade to Starter or Pro to continue testing without losing any work.
Can I pay internationally from outside India with a global card?keyboard_arrow_down
Yes. Tanvelo's checkout supports international payments in USD ($) and INR (₹). We accept all major international credit and debit cards—including Visa, Mastercard, American Express, Diners Club, and international corporate cards—from over 130 countries worldwide.
What is the difference between '4 core' and 'All 10' categories?keyboard_arrow_down
The Trial tier includes our 4 foundational categories: Reasoning Rigor, Knowledge Grounding, Deterministic Coding, and Safety Separation. Starter unlocks all 10 domain categories—including Instruction Drift, Multi-turn Context Retention, Semantic Hallucination Stress, Jailbreak Resilience, Edge Syntax, and Custom Business Requirements. Pro further adds bulk automated suite generation and custom rubric criteria.
How does the 'Max per run' limit work?keyboard_arrow_down
Max per run defines the maximum number of stress-test prompts generated and judged within a single benchmark session. On Trial you can test up to 50 prompts per run; on Starter up to 300; and on Pro up to 500 prompts per run. This allows deeper statistical confidence and fine-grained cluster resolution on complex models.
How does 'compare 2 models' work on Starter and Pro?keyboard_arrow_down
On Starter and Pro, Tanvelo provides side-by-side regression analysis. You can benchmark release candidate v2.1 against your baseline v2.0 on the exact same stress suite to visually compare score deltas, verify resolved failure clusters, and ensure no regressions were introduced before production deployment.
What format do exported remediation datasets come in?keyboard_arrow_down
On Starter and Pro, datasets are exported in standard JSONL (JSON Lines) and CSV formats. Each entry includes the original failure prompt, the categorized failure mode, and a verified gold-standard target completion designed directly for supervised fine-tuning (SFT) or prompt regression suites. Pro also adds audit-ready PDF executive summaries.
Are my API keys and model completions private?keyboard_arrow_down
Yes, 100%. All API credentials are encrypted with AES-256 (Fernet) at rest. In accordance with PRD Section 43, Tanvelo never uses your inputs, completions, or diagnostic results to train foundational models. Your evaluations remain completely private to your workspace.
Begin diagnosing your AI model in under 2 minutes.
Start free with 50 automated tests, LLM judge evaluation, safety screening, and top 3 failure clusters. No credit card required.
*Trial ends on first limit hit. Upgrade to continue.