Forge Reliability Scoreboard

Community-generated AI reliability benchmarks — cryptographically signed, tamper-evident.

Not a vendor benchmark. Not self-reported. Real inference. Real signatures.

5
Models Tested
24
Signed Runs
8
Scenario Categories
Ed25519
Signature Algorithm
90–100%
Excellent — strong adversarial reliability
75–89%
Good — reliable with minor gaps
<75%
Needs work — significant weaknesses

Reliability Rankings

# Model Avg Score Best Score Runs Scenarios Latest Report
1 Qwen/Qwen2.5-72B-Instruct-AWQ 72.7% 72.7% 2 161 f6a343bd346c…
2 Qwen/Qwen2.5-32B-Instruct-AWQ 65.2% 65.2% 2 161 670ff8d0dcee…
3 Qwen/Qwen2.5-3B-Instruct-AWQ 63.4% 63.4% 2 161 7b5831135092…
4 stelterlab/Mistral-Small-24B-Instruct-2501-AWQ 63% 63.4% 2 161 cf12ce34e3eb…
5 LGAI-EXAONE/EXAONE-3.5-7.8B-Instruct-AWQ 57.1% 57.1% 16 161 3a07e1f15480…

Average score is computed across all signed runs for that model. Each run is verified against the submitter's Ed25519 machine key.

Recent Reports

Run IDModelScoreSubmitted
3a07e1f1548021f4 LGAI-EXAONE/EXAONE-3.5-7.8B-Instruct-AWQ 57.1% 2026-07-02 15:04
64f1a5a5581151d8 LGAI-EXAONE/EXAONE-3.5-7.8B-Instruct-AWQ 57.1% 2026-07-02 14:13
ec0fa5f8ca327a95 LGAI-EXAONE/EXAONE-3.5-7.8B-Instruct-AWQ 57.1% 2026-07-01 10:10
91d0d935933e6b3e LGAI-EXAONE/EXAONE-3.5-7.8B-Instruct-AWQ 57.1% 2026-07-01 09:22
56c08f284322c217 LGAI-EXAONE/EXAONE-3.5-7.8B-Instruct-AWQ 57.1% 2026-06-29 03:50
f9e8741cffa064d3 LGAI-EXAONE/EXAONE-3.5-7.8B-Instruct-AWQ 57.1% 2026-06-29 03:03
d520fa33ab69f131 LGAI-EXAONE/EXAONE-3.5-7.8B-Instruct-AWQ 57.1% 2026-06-28 22:44
f91fd7a2a356d19e LGAI-EXAONE/EXAONE-3.5-7.8B-Instruct-AWQ 57.1% 2026-06-28 21:59
130ba69e04612b4e LGAI-EXAONE/EXAONE-3.5-7.8B-Instruct-AWQ 57.1% 2026-06-28 16:04
95b18fcaff55cc1e LGAI-EXAONE/EXAONE-3.5-7.8B-Instruct-AWQ 57.1% 2026-06-28 15:11
0f5a136bcb2e126e LGAI-EXAONE/EXAONE-3.5-7.8B-Instruct-AWQ 57.1% 2026-06-27 19:03
0a3e53e8d163d935 LGAI-EXAONE/EXAONE-3.5-7.8B-Instruct-AWQ 57.1% 2026-06-27 18:15
9e3b7489074798c7 LGAI-EXAONE/EXAONE-3.5-7.8B-Instruct-AWQ 57.1% 2026-06-25 21:49
58d2337511e1b48c LGAI-EXAONE/EXAONE-3.5-7.8B-Instruct-AWQ 57.1% 2026-06-25 20:04
670ff8d0dcee962e Qwen/Qwen2.5-32B-Instruct-AWQ 65.2% 2026-06-24 23:57
fb1f2eb1fe7cb80c Qwen/Qwen2.5-32B-Instruct-AWQ 65.2% 2026-06-24 23:05
f6a343bd346caea6 Qwen/Qwen2.5-72B-Instruct-AWQ 72.7% 2026-06-23 14:54
ae33b1122020c0e7 Qwen/Qwen2.5-72B-Instruct-AWQ 72.7% 2026-06-23 14:11
1b23f59b297d3b8b LGAI-EXAONE/EXAONE-3.5-7.8B-Instruct-AWQ 57.1% 2026-06-23 01:47
20827ccb462dcb51 LGAI-EXAONE/EXAONE-3.5-7.8B-Instruct-AWQ 57.1% 2026-06-23 00:56

Add your model to the scoreboard

forge break --model qwen3:14b --share

Requires a Forge BPoS passport. Results are signed with your machine's Ed25519 key and verified server-side before appearing. Forge is model-agnostic — local models, OpenAI, Anthropic, and any Ollama-compatible backend are all supported.

GitHub  ·  JSON API