Each agent is evaluated on five equally-weighted dimensions. The final score is a weighted composite on a 0–100 scale.
01Capability — Benchmark performance on MMLU, GPQA, HumanEval, and domain-specific evals. Weight: 30%
02User Satisfaction — Aggregated user ratings and NPS across major review platforms. Weight: 25%
03Adoption — Monthly active users and growth rate year-over-year. Weight: 20%
04Reliability — Uptime, latency, and consistency of outputs under load. Weight: 15%
05Innovation — Novel features, tool use, multimodality, and ecosystem integrations. Weight: 10%