Compare/DeepHermes 3 - Llama-3.1 8B Preview (Non-reasoning) vs GLM-4.5V (Non-reasoning)

DeepHermes 3 - Llama-3.1 8B Preview (Non-reasoning)vsGLM-4.5V (Non-reasoning)

Side-by-side comparison of pricing, 12 benchmarks, and generation speed.

Nous Research

DeepHermes 3 - Llama-3.1 8B Preview (Non-reasoning)

Input
Output
Speed
TTFT
Z AI

GLM-4.5V (Non-reasoning)

Input
$0.6/M
Output
$1.8/M
Speed
92 tok/s
TTFT
1.85s

Winner by Category

Cheaper
GLM-4.5V (Non-reasoning)
Faster (tok/s)
GLM-4.5V (Non-reasoning)
Lower Latency
GLM-4.5V (Non-reasoning)
Benchmarks (0-1)
GLM-4.5V (Non-reasoning)

Pricing Comparison

MetricDeepHermes 3 - Llama-3.1 8B Preview (Non-reasoning)GLM-4.5V (Non-reasoning)
Input ($/M tokens)$0.6
Output ($/M tokens)$1.8
Cost for 1M input + 100K output tokens:
GLM-4.5V (Non-reasoning)$0.78

Speed Comparison

Output Speed (tokens/s) — higher is better
DeepHermes 3 - Llama-3.1 8B Preview (Non-reasoning)
GLM-4.5V (Non-reasoning)
92 tok/s
Time to First Token (seconds) — lower is better
DeepHermes 3 - Llama-3.1 8B Preview (Non-reasoning)
GLM-4.5V (Non-reasoning)
1.85s

Editorial Analysis

Verdict. GLM-4.5V (Non-reasoning) takes the aggregate benchmark matchup 1–0 across 1 categories. Real workloads usually care about a handful of specific tasks — see the per-benchmark table above.

Pricing. Pricing varies significantly between these models — check the table above for the exact per-token rates. Many production workloads actually surface input-token cost (retrieval-augmented prompts, code-context windows), so factor both directions.

Strengths. DeepHermes 3 - Llama-3.1 8B Preview (Non-reasoning) is strongest on Intelligence Index (1.9). GLM-4.5V (Non-reasoning) leads on Intelligence Index (6.8).

Speed. Speed data is incomplete for this pair; benchmark and price should decide.

Provider. Nous Research and Z AI sell to overlapping but distinct developer audiences: Nous Research tends to ship frontier reasoning models with premium positioning, while Z AI often prices more aggressively. Your existing vendor relationships, billing, and SLA preferences may matter as much as the raw numbers above.

Recommendation. Both models have legitimate use cases — the right answer depends on whether you are optimizing for benchmark ceiling, latency, or unit cost. Start with the cheaper / faster model, evaluate against your specific task, and only switch if the upgrade shows a meaningful lift.

Benchmark Comparison

Data from Artificial Analysis API — 12 benchmarks

Intelligence Index
1.96.8
Coding Index
Math Index
GPQA Diamond
MMLU-Pro
LiveCodeBench
AIME 2025
MATH-500
Humanity's Last Exam
SciCode
IFBench
TerminalBench
DeepHermes 3 - Llama-3.1 8B Preview (Non-reasoning)0 wins
1 winsGLM-4.5V (Non-reasoning)

Frequently Asked Questions

Which is cheaper, DeepHermes 3 - Llama-3.1 8B Preview (Non-reasoning) or GLM-4.5V (Non-reasoning)?

GLM-4.5V (Non-reasoning) is cheaper overall. Its blended price (3:1 input/output ratio) is $0.90/M tokens vs $—/M for DeepHermes 3 - Llama-3.1 8B Preview (Non-reasoning).

Which model performs better on benchmarks?

GLM-4.5V (Non-reasoning) wins 1 out of 12 benchmarks compared to 0 for DeepHermes 3 - Llama-3.1 8B Preview (Non-reasoning). See the detailed benchmark chart above for per-category results.

Which is faster for real-time applications?

GLM-4.5V (Non-reasoning) generates tokens faster at 92 tok/s vs — tok/s. However, GLM-4.5V (Non-reasoning) has lower time-to-first-token (1.85s vs —s).

When should I use DeepHermes 3 - Llama-3.1 8B Preview (Non-reasoning) vs GLM-4.5V (Non-reasoning)?

Choose based on your priorities: GLM-4.5V (Non-reasoning) for lower cost, GLM-4.5V (Non-reasoning) for stronger benchmark performance, and GLM-4.5V (Non-reasoning) for faster generation. For latency-sensitive apps, check the TTFT comparison above.