Side-by-side comparison of pricing, 12 benchmarks, and generation speed.
| Metric | Gemini 2.0 Flash-Lite (Preview) | GLM-4.5V (Non-reasoning) |
|---|---|---|
| Input ($/M tokens) | — | $0.6 |
| Output ($/M tokens) | — | $1.8 |
Verdict. Gemini 2.0 Flash-Lite (Preview) wins the overall benchmark matchup 1–0 across 1 overlapping categories, but raw benchmark score is only one input to the decision.
Pricing. Pricing varies significantly between these models — check the table above for the exact per-token rates. Many production workloads actually surface input-token cost (retrieval-augmented prompts, code-context windows), so factor both directions.
Strengths. Gemini 2.0 Flash-Lite (Preview) is strongest on Intelligence Index (8.4). GLM-4.5V (Non-reasoning) leads on Intelligence Index (6.8).
Speed. Speed data is incomplete for this pair; benchmark and price should decide.
Provider. Google and Z AI sell to overlapping but distinct developer audiences: Google tends to ship frontier reasoning models with premium positioning, while Z AI often prices more aggressively. Your existing vendor relationships, billing, and SLA preferences may matter as much as the raw numbers above.
Recommendation. Both models have legitimate use cases — the right answer depends on whether you are optimizing for benchmark ceiling, latency, or unit cost. Start with the cheaper / faster model, evaluate against your specific task, and only switch if the upgrade shows a meaningful lift.
Data from Artificial Analysis API — 12 benchmarks
GLM-4.5V (Non-reasoning) is cheaper overall. Its blended price (3:1 input/output ratio) is $0.90/M tokens vs $—/M for Gemini 2.0 Flash-Lite (Preview).
Gemini 2.0 Flash-Lite (Preview) wins 1 out of 12 benchmarks compared to 0 for GLM-4.5V (Non-reasoning). See the detailed benchmark chart above for per-category results.
GLM-4.5V (Non-reasoning) generates tokens faster at 92 tok/s vs — tok/s. However, GLM-4.5V (Non-reasoning) has lower time-to-first-token (1.85s vs —s).
Choose based on your priorities: GLM-4.5V (Non-reasoning) for lower cost, Gemini 2.0 Flash-Lite (Preview) for stronger benchmark performance, and GLM-4.5V (Non-reasoning) for faster generation. For latency-sensitive apps, check the TTFT comparison above.