DEV Community

#benchmarking

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
Don't Trust a New Model's Benchmarks Until You Run Your Own 30-Minute Smoke Test

Don't Trust a New Model's Benchmarks Until You Run Your Own 30-Minute Smoke Test

Comments
5 min read
Before You Adopt MiniMax H3, Run a Twenty-Minute Model Audit

Before You Adopt MiniMax H3, Run a Twenty-Minute Model Audit

Comments
3 min read
DeepSeek Harness: What "Everything is a Plugin" Actually Means for Agent Frameworks

DeepSeek Harness: What "Everything is a Plugin" Actually Means for Agent Frameworks

Comments
2 min read
Free Model Endpoints Are an Evaluation Problem, Not a Hosting Strategy

Free Model Endpoints Are an Evaluation Problem, Not a Hosting Strategy

Comments
3 min read
New Benchmark for Evaluating Long-Horizon Agents in Online Environments

New Benchmark for Evaluating Long-Horizon Agents in Online Environments

Comments
4 min read
Native HTTP Engine for Node: Performance Benchmarks Against uWS, Bun, Fastify, and Hono

Native HTTP Engine for Node: Performance Benchmarks Against uWS, Bun, Fastify, and Hono

Comments
3 min read
Microsoft measured our thesis, and we still cannot quote ours

Microsoft measured our thesis, and we still cannot quote ours

1
Comments
7 min read
Where to split a sentence for streaming TTS is decided by one number

Where to split a sentence for streaming TTS is decided by one number

Comments
5 min read
Measuring LLM Prefix Caching: The Cache Hit Rate Metric

Measuring LLM Prefix Caching: The Cache Hit Rate Metric

Comments
6 min read
llmperf Is Archived: Alternatives for LLM Benchmarking

llmperf Is Archived: Alternatives for LLM Benchmarking

Comments
3 min read
My context selector beat grep. An agent with grep beat it.

My context selector beat grep. An agent with grep beat it.

Comments
7 min read
Correctness Has a Price: We Benchmarked Fair Leaderboards