Skip to content
Navigation menu
Search
Powered by Algolia
Search
Log in
Create account
DEV Community
Close
#
benchmarking
Follow
Hide
Posts
Left menu
đź‘‹
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
Right menu
Don't Trust a New Model's Benchmarks Until You Run Your Own 30-Minute Smoke Test
Dakota Wu
Dakota Wu
Dakota Wu
Follow
Aug 17
Don't Trust a New Model's Benchmarks Until You Run Your Own 30-Minute Smoke Test
#
ai
#
programming
#
benchmarking
#
opensource
Comments
Add Comment
5 min read
Before You Adopt MiniMax H3, Run a Twenty-Minute Model Audit
Casey Li
Casey Li
Casey Li
Follow
Aug 14
Before You Adopt MiniMax H3, Run a Twenty-Minute Model Audit
#
ai
#
opensource
#
programming
#
benchmarking
Comments
Add Comment
3 min read
DeepSeek Harness: What "Everything is a Plugin" Actually Means for Agent Frameworks
Cole Halton
Cole Halton
Cole Halton
Follow
Aug 13
DeepSeek Harness: What "Everything is a Plugin" Actually Means for Agent Frameworks
#
deepseek
#
aiagents
#
opensource
#
benchmarking
Comments
Add Comment
2 min read
Free Model Endpoints Are an Evaluation Problem, Not a Hosting Strategy
Riley Xu
Riley Xu
Riley Xu
Follow
Aug 14
Free Model Endpoints Are an Evaluation Problem, Not a Hosting Strategy
#
ai
#
opensource
#
benchmarking
#
python
Comments
Add Comment
3 min read
New Benchmark for Evaluating Long-Horizon Agents in Online Environments
David DĂaz
David DĂaz
David DĂaz
Follow
Aug 12
New Benchmark for Evaluating Long-Horizon Agents in Online Environments
#
ai
#
benchmarking
#
agents
#
onlineservices
Comments
Add Comment
4 min read
Native HTTP Engine for Node: Performance Benchmarks Against uWS, Bun, Fastify, and Hono
Raiyan C
Raiyan C
Raiyan C
Follow
Aug 10
Native HTTP Engine for Node: Performance Benchmarks Against uWS, Bun, Fastify, and Hono
#
node
#
performance
#
http
#
benchmarking
Comments
Add Comment
3 min read
Microsoft measured our thesis, and we still cannot quote ours
Tom Jones
Tom Jones
Tom Jones
Follow
Aug 9
Microsoft measured our thesis, and we still cannot quote ours
#
ai
#
testing
#
benchmarking
#
programming
1
 reaction
Comments
Add Comment
7 min read
Where to split a sentence for streaming TTS is decided by one number
Renga
Renga
Renga
Follow
Aug 11
Where to split a sentence for streaming TTS is decided by one number
#
performance
#
typescript
#
ai
#
benchmarking
Comments
Add Comment
5 min read
Measuring LLM Prefix Caching: The Cache Hit Rate Metric
Wayne
Wayne
Wayne
Follow
Aug 5
Measuring LLM Prefix Caching: The Cache Hit Rate Metric
#
llm
#
benchmarking
#
performance
Comments
Add Comment
6 min read
llmperf Is Archived: Alternatives for LLM Benchmarking
Wayne
Wayne
Wayne
Follow
Aug 5
llmperf Is Archived: Alternatives for LLM Benchmarking
#
rust
#
benchmarking
#
llm
Comments
Add Comment
3 min read
My context selector beat grep. An agent with grep beat it.
Anish Shrestha
Anish Shrestha
Anish Shrestha
Follow
Jul 31
My context selector beat grep. An agent with grep beat it.
#
contextengineering
#
retrieval
#
swebench
#
benchmarking
Comments
Add Comment
7 min read
Correctness Has a Price: We Benchmarked Fair Leaderboards
Trung Duong
Trung Duong