302.AI Model Leaderboard
Quickly calculate costs, screen high-value models, compare multi-model capabilities, precise selection, and one-click try-out
50+ Hot Models | 17+ Brands | 6+ Evaluation Dimensions
Updated At:2026-09-15
Cost-Performance Leaderboard
Multi-dimensional comparison & computation (Comprehensive Capability, Market Price, Blended Price) using official ranking data of AA Intelligence, ECI and LM Arena. All model values take qwen3.7-max as the benchmark. Higher CP score means better cost-effectiveness.
Calculation Method of Cost-Performance (CP)
Expand MoreCost-performance analysis samples mainstream LLMs with rankings from AA Intelligence, ECI, LM Arena. Calculations adopt three dimensions: comprehensive capability, input-output weighted blended price and LOWESS-fitted market price. All models are normalized against qwen3.7-max
Comprehensive Capability score (I) is weighted: AA Intelligence (40%) + normalized ECI score (30%) + normalized LM Arena score (30%). Cross-verified by multiple authoritative leaderboards, higher Comprehensive Capability score means stronger model
Formula :
Perform LOWESS locally weighted regression fitting of based on the full pool samples (I, P) with parameters (frac=0.8, it=2). The benchmark market curve is generated via exponential restoration using . When a model has a comprehensive capability score I, the theoretically fair market reference price fitted from all models in the pool is defined as the market price
Formula :
Blended Price is calculated from input price (75%) and output price (25%) to facilitate comparison of the overall pricing across all models
Formula :
CP = ( 3.087789 / 2.675 ) / ( 3.087789 / 2.675 ) × 60 = 60CP = ( 1.772616 / 0.325 ) / ( 3.087789 / 2.675 ) × 60 = 283.5Benchmark Model
Benchmark model is qwen3.7-max, with its cost-performance index fixed at 60. This model acts as the comparison benchmark to conduct lateral comparative analysis with alternative models
Parameter Definitions: Io represents the comprehensive capability score of benchmark model qwen3.7-max; Po represents the blended price of benchmark model qwen3.7-max
Benchmark Market Line (green dashed line)
All (I, P) samples are fitted to with LOWESS (parameters: frac=0.8, it=2). Exponentiate the fitting outcome with to get the Benchmark Market Line, indicating the benchmark market price matching each comprehensive capability level
Cost-Effective Zone (green area)
The area to the left of the Benchmark Market Line is defined as the cost-effective zone. If a model's scatter point lies within this zone, its actual blended price is lower than the market price at the same comprehensive capability level, signifying superior cost performance
Normalization
Linear normalization is a data preprocessing method that linearly maps the original index values proportionally to a standardized range. In this instance, linear normalization is applied to the original scores of ECI and LM Arena. Original value ranges for each metric are listed below: LM Arena: [800, 1800], ECI: [45, 200]
Formula:
Overall Leaderboard
Comprehensive Capability score is calculated via weighted aggregation: AA Intelligence (40%), normalized ECI score (30%), and normalized LM Arena score (30%), with cross-verification across multiple authoritative lists. Higher comprehensive capability scores correspond to stronger overall model performance
Brand manufacturer filtering is supported to intuitively compare differences among various models under the same brand, enabling horizontal selection for internal variants such as the Claude and GPT series
Response Speed Leaderboard
Time To First Token (TTFT) serves as a standardized measurement metric. Lower values indicate faster model response times
Response speed is prioritized for real-time conversation, streaming output and high-frequency invocation scenarios. Low-latency models deliver better user experience than high-score yet slow-response models
Coding Leaderboard
Calculated by weighted combination of normalized ECI Software Engineering score (60%) and normalized LM Arena Coding score (40%), verified crosswise via multiple authoritative leaderboards. Higher coding capability score indicates superior model performance, making it the optimal choice for programming, bug fixing and vibe coding scenarios
DeepSeek R1 balances reasoning and coding capabilities, offering excellent cost performance. The Claude series excels in delivering the most well-structured and standardized code
Multimodal Leaderboard
Calculated through weighted computation of normalized LM Arena Vision score (60%) and AA Multimodal score (40%), cross-verified by multiple authoritative leaderboards. Higher multimodal scores correspond to better performance in image recognition, image-text comprehension and visual question answering tasks
Supports screenshot recognition, chart analysis, and scanned document reading. Simply upload an image and ask questions directly without textual description of content
Vision Leaderboard
Rankings standardized by real tests from 302.AI Benchmark Lab. Higher rankings indicate stronger overall image and audio-video processing performance


























