r/LocalLLaMA · · 1 min read

Do you think a few Qwen3.8-27B models working together could score as well as Fable-5 on LiveCodeBench Hard?

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

Do you think a few Qwen3.8-27B models working together could score as well as Fable-5 on LiveCodeBench Hard?

Has anyone tested this? Ensemble of small Qwen models claiming Fable 5-level coding performance.

A new paper claims that running several Qwen3.8-27B models together matches Fable 5’s accuracy on LiveCodeBench.

The authors also say their setup paired with GPT Terra reaches Fable 5-level coding accuracy on LiveCodeBench at roughly a fifth of the cost.

Curious what people here think; is this worth actually trying out?

https://github.com/slee-persis/GVS5H

https://arxiv.org/abs/2608.26480

submitted by /u/sl4447
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA