AgentLens
GitHub
AI OBSERVABILITY + LLM MONITORING

Choose the right AI.Monitor it in production.Fix failures before users notice.

Compare AI providers before launch, then monitor OpenAI, Anthropic, Gemini, Groq and other LLMs in production. AgentLens explains reliability, latency, errors and cost without collecting your prompts.

Explore the source
Transparent scoring Provider keys stay local One-command tests
agentlens / production
LIVE
INCIDENT 412

Provider degradation

Last 60 min
Success rate61.2%↓ 37.2%
p95 latency2.8s3.1× slower
Calls observed12,402live sample
baseline 98.4%
12:0012:2012:4013:00
No prompt content collectedUpdated 4s ago
Likely root causeProvider rate limitingConfidence 94%
Safest route nowSwitch to Anthropic99.1% success · 410ms
OPENAIANTHROPICGROQGEMINIMISTRALAZURE AIOPENAIANTHROPICGROQGEMINIMISTRALAZURE AI
7supported provider paths
3visible synthetic checks
1local runner command
0raw responses uploaded
FROM SIGNAL TO ANSWER

Not another dashboard.
A decision, with evidence.

AgentLens turns scattered provider behavior into one clear incident story: what changed, why it matters, and what your team should do next.

Baseline normal

Latency begins rising

Success rate drops

AgentLens opens incident

Safer route recommended

AGENTLENS ANSWERReady

Route new traffic to Anthropic.

OpenAI is rate limiting this workload. Anthropic is handling a matching request profile at 99.1% success with normal latency.

94% root-cause confidence12,402 calls compared
THE DIFFERENCE

Your support inbox should never be your monitoring system.

WITHOUT AGENTLENS
Customer complaintProvider dashboardRaw logsGuess the cause
45 minutes laterStill investigating.
WITH AGENTLENS
Incident detected

OpenAI success fell to 61%.

Root causeRate limiting

Safest actionRoute to Anthropic

Answer in under 60 secondsWith evidence attached.
HOW IT WORKS

From idea to evidence.
Then production confidence.

One guided path for choosing, connecting, and monitoring AI—without replacing the stack you already use.

01

Describe

Explain the AI job in normal language. No benchmark knowledge needed.

02

Compare

Run the same synthetic checks locally across your selected providers.

03

Choose

See the best fit for your quality, speed, reliability, and cost priorities.

04

Connect

Install one helper without replacing your existing AI integration.

05

Monitor

Keep watching the choice with evidence from real production behavior.

Guided AgentLens setupnpx agentlens init
BUILT FOR THE INCIDENT

Every view moves the
investigation forward.

Hover and explore: each surface reveals the evidence behind the recommendation instead of hiding it behind decoration.

01 / INCIDENT INTELLIGENCE

A timeline that reconstructs what happened.

Changes are ordered, correlated, and attached to the provider, model, release, and route that produced them.

Explore an incident
12:41Baseline
12:43Latency
12:45Retries
12:47Detected
12:49Alerted
02 / AUTOMATIC ALERTS

Context arrives where your team works.

No “something is wrong” alerts. Every notification carries the signal, cause, and next step.

#production-aiOpenAI success fell to 61%.

now

Issue #412 createdInvestigation context attached.

4s

Webhook delivered200 OK · 84ms

5s
03 / INVESTIGATION PROMPT

Give your engineer a useful starting point.

AgentLens packages the important evidence into a ready-to-use investigation prompt.

investigate.md

Investigate a production degradation affecting OpenAI / gpt-4o.

Observed: success rate fell from 98.4% to 61.2%...

PROVIDER MATCH

No universal “best model.”
Only the best fit for your job.

Every provider receives the same visible synthetic checks. AgentLens weights instruction fit, reliability, speed, and cost around what matters to your product.

ProviderReliabilityLatencyDecision
OpenAI
61%
2.8sDegraded
Anthropic
99%
410msRecommended
Groq
100%
180msFastest
Gemini
97%
520msStable
EXPLAINABLE BY DEFAULT

The recommendation is only useful if you can trust it.

Every answer links back to the observations that produced it. No opaque reliability score and no hidden prompt content.

  • Production telemetry only
  • Evidence attached to every conclusion
  • Provider-neutral recommendations
INVESTIGATION / 41294% confidence

Why did success rate fall?

1

Request volumeNormal · not causal

2

Application releaseNo changes in window

3

Provider response codes429 responses increased 8.6×

4

ConclusionRate limiting is the likely cause

OPEN SOURCE

Your reliability layer
should not be a black box.

Inspect the collection path, run the stack yourself, and keep ownership of your operational data.

TypeScript Python Self-hostable Multi-provider
Explore AgentLens on GitHub
SIMPLE PRICING

Test before you commit.
Pay when you grow.

Clear monthly pricing in INR. Provider usage is paid directly through your own accounts; AgentLens checkout is handled by Razorpay.

FREE

₹0

For choosing and watching your first AI feature.

  • 3 comparisons each month
  • 1 monitored project
  • Local private runner
SCALE

₹4,999 /month

For teams evaluating AI across multiple products.

  • 200 comparisons each month
  • Up to 25 monitored projects
  • Priority setup support
SEE IT BEFORE THEY DO

Your next AI incident
already has an answer.

Start monitoring real production behavior in minutes.

View source