Transparent Tiering

Calibrated plans for deterministic AI.

From exploratory model stress testing to production failure clustering and fine-tuning remediation datasets. Zero hidden surcharges.

publicInternational Payments Enabled (130+ countries)
Standard monthly billing • Cancel or upgrade anytime
Evaluation5-day duration

Trial

Experience Tanvelo's diagnostic architecture and semantic failure clustering on your first model.

$0/ free
Test Capacity:50 tests / 5 days
Concurrency:Max 50 / run
Workspace:1 project • 1 model • 1 seat
check_circle4 core evaluation categories
check_circleLLM judge + Safety evaluation
check_circleScore + top 3 failure clusters
check_circlePreview remediation datasets
check_circleDashboard export access
check_circle5 days evaluation history
check_circleStandard execution queue
Start freearrow_forward

*Trial ends on first limit hit. Upgrade to continue.

starMost Popular
Continuous DevPriority Queue

Starter

For active engineering teams running frequent model stress tests, comparisons, and root-cause remediation.

$20/mo
Monthly Tests:5,000 / month
Max per run:300 test cases
Workspace:3 projects • 2 models/project
check_circleAll 10 categories + custom requirements
check_circleFull evaluation pipeline & rubric scoring
check_circleFull root-cause + actionable recommendations
check_circle500 / month targeted remediation datasets
check_circleJSON + CSV exports + compare 2 models
check_circle90 days evaluation history retention
check_circlePriority queue dispatch for faster test execution

Pay with Razorpay • Cards, UPI & International

Production Teams5 Seats • Audit

Pro

For high-velocity AI teams demanding unlimited projects, multi-seat collaboration, and audit-ready reports.

$50/mo
Monthly Tests:10,000 / month
Max per run:500 test cases
Workspace:Unlimited projects • 5 seats
check_circleAll categories + bulk generation + custom criteria
check_circleFull evaluation pipeline
check_circleFull + prioritized root-cause diagnosis
check_circle2,000 / month remediation datasets
check_circleJSON + CSV + PDF reports + audit log export
check_circle12 months evaluation history retention
check_circleFastest queue (small runs first priority)

Pay with Razorpay • Cards, UPI & International

credit_card
Worldwide International Payments Supported via Razorpay130+ Countries

Pay securely in USD ($) or INR (₹) using Visa, Mastercard, American Express, Apple Pay, Discover, Diners Club, or domestic UPI.

verified_user256-bit Encrypted
info

*Trial ends on first limit hit. Upgrade to continue.

When you hit 50 tests or reach 5 days, benchmark execution pauses. All historical logs, reports, and clusters are preserved indefinitely.

Start 5-day trial →
Comprehensive Matrix

Side-by-side tier specification

Compare diagnostic limits, evaluation depth, export rights, and queue priorities across every plan.

Viewing: starter Plan ($20)

Use the toggle buttons above to switch plans

tuneLimits & Capacity
Tests
Total automated diagnostic test evaluations executed per billing period.
5,000 / month
Projects
Independent application or workspace containers to isolate models and datasets.
3
Models
Connected LLM endpoints (OpenAI, Ollama, vLLM, OpenRouter, or custom APIs) per project.
2 per project
Seats
Team members with synchronized workspace access and shared diagnostic reports.
1
Max per run
Maximum parallel benchmark test cases executed in a single evaluation run.
300
psychologyBenchmarking & Evaluation
Categories
Diagnostic failure domains: Reasoning, Grounding, Coding, Safety, plus Drift, Jailbreak, Edge Cases, etc.
All 10 + custom requirements
Evaluation
Decoupled judging architecture with calibrated rubrics and automated safety checks.
Full evaluation
Diagnosis
Semantic clustering of model failures into actionable root causes.
Full root-cause + recommendations
datasetDatasets, Exports & Retention
Datasets
Supervised fine-tuning pairs (input prompt + verified gold completion) ready to download.
500 / mo
Exports
Available export formats and model regression comparison capabilities.
JSON + CSV + compare 2 models
History
Duration for which full benchmark logs, failure clusters, and generated traces are retained.
90 days
speedExecution Speed & Queue
Speed
Runner queue dispatch priority and concurrent worker allocation.
Priority queue
apartmentEnterprise & Private Deployments

Need custom VPC runners, private judge rubrics, or 100K+ monthly tests?

Tanvelo offers dedicated isolated test runners inside your cloud perimeter (AWS, GCP, Azure), custom domain-specific synthetic generation pipelines, custom evaluation judges, and SLAs for mission-critical deployments.

Plan Inquiries

Frequently Asked Questions

Clear terms and explanations regarding quotas, limits, international cards, and evaluations.

What does 'Trial ends on first limit hit. Upgrade to continue' mean?keyboard_arrow_down

The Trial is completely free ($0) with no credit card required. You receive 50 tests or 5 days of access (whichever comes first) on 1 connected model and 1 project. Once you exhaust 50 tests or the 5 days expire, benchmark execution pauses. All your historical logs, score cards, and clusters remain safely preserved, and you can instantly upgrade to Starter or Pro to continue testing without losing any work.

Can I pay internationally from outside India with a global card?keyboard_arrow_down

Yes. Tanvelo's checkout supports international payments in USD ($) and INR (₹). We accept all major international credit and debit cards—including Visa, Mastercard, American Express, Diners Club, and international corporate cards—from over 130 countries worldwide.

What is the difference between '4 core' and 'All 10' categories?keyboard_arrow_down

The Trial tier includes our 4 foundational categories: Reasoning Rigor, Knowledge Grounding, Deterministic Coding, and Safety Separation. Starter unlocks all 10 domain categories—including Instruction Drift, Multi-turn Context Retention, Semantic Hallucination Stress, Jailbreak Resilience, Edge Syntax, and Custom Business Requirements. Pro further adds bulk automated suite generation and custom rubric criteria.

How does the 'Max per run' limit work?keyboard_arrow_down

Max per run defines the maximum number of stress-test prompts generated and judged within a single benchmark session. On Trial you can test up to 50 prompts per run; on Starter up to 300; and on Pro up to 500 prompts per run. This allows deeper statistical confidence and fine-grained cluster resolution on complex models.

How does 'compare 2 models' work on Starter and Pro?keyboard_arrow_down

On Starter and Pro, Tanvelo provides side-by-side regression analysis. You can benchmark release candidate v2.1 against your baseline v2.0 on the exact same stress suite to visually compare score deltas, verify resolved failure clusters, and ensure no regressions were introduced before production deployment.

What format do exported remediation datasets come in?keyboard_arrow_down

On Starter and Pro, datasets are exported in standard JSONL (JSON Lines) and CSV formats. Each entry includes the original failure prompt, the categorized failure mode, and a verified gold-standard target completion designed directly for supervised fine-tuning (SFT) or prompt regression suites. Pro also adds audit-ready PDF executive summaries.

Are my API keys and model completions private?keyboard_arrow_down

Yes, 100%. All API credentials are encrypted with AES-256 (Fernet) at rest. In accordance with PRD Section 43, Tanvelo never uses your inputs, completions, or diagnostic results to train foundational models. Your evaluations remain completely private to your workspace.

Zero Obligation Benchmark

Begin diagnosing your AI model in under 2 minutes.

Start free with 50 automated tests, LLM judge evaluation, safety screening, and top 3 failure clusters. No credit card required.

*Trial ends on first limit hit. Upgrade to continue.