🚩 Report: Spam

#7
by DedeProGames - opened

Subject: Leaderboard fairness: maj@16 self-reported score ranked #1 alongside single-sample scores (FINAL-Bench/Darwin-180B-RSI)

Hello Hugging Face team,

I'd like to flag a possible fairness issue in the official benchmark leaderboards.

Model: https://huggingface.co/FINAL-Bench/Darwin-180B-RSI

What I observed

  • The model card claims #1 on the GPQA Diamond, MMLU-Pro and MMMU-Pro leaderboards.
  • The GPQA Diamond score (94.44) is majority vote over up to 16 samples, with a 131,072-token thinking budget. The MMMU-Pro score (79.48) is majority vote over 3 samples. Only the MMLU-Pro score (88.12) is single-sample.
  • All three are labeled self-reported in the eval results, and the card says "All numbers are self-measured."
  • The leaderboard appears to rank these entries together with other models' scores without distinguishing the protocol (single-sample vs. multi-sample voting).

Why this matters

  • Multi-sample majority voting multiplies inference compute, and it isn't comparable to single-sample results. GPQA Diamond has only 198 questions, so each question is worth ~0.5 points.
  • The publisher confirmed in the discussion that the single-sample score for a sibling model (Darwin-27B-Opus) is 72.85, while the card's family table lists 86.9 for it. This shows how much the protocol changes the result.
  • The card doesn't report the parent model (Qwen3.8-Flash-Next) under the same maj@16 protocol, so the effect of the claimed training method can't be separated from the effect of voting. The only same-protocol comparison shown (MMLU-Pro, 88.04 vs 88.12) is within noise.
  • No evaluation harness or per-question outputs for this model are linked, so the numbers can't be independently re-scored.

Discussion for context: https://huggingface.co/FINAL-Bench/Darwin-180B-RSI/discussions/1 (closed by the publisher without answering questions about the parent's score under the same protocol, single-sample results, or the release of the harness and outputs).

Request
Could you review whether these entries should be ranked as #1 alongside single-sample results? For example, by separating leaderboards or columns by protocol (single-sample vs. maj@k), marking self-reported entries distinctly, or requiring reproducible evaluation logs for top positions.

Thank you for your time.

FINAL_Bench org

For the Hugging Face team:

The author of this thread has repeatedly posted insulting and mocking comments on our model's discussion page (e.g. "are you stupid?").

Every score we submitted states its protocol (number of samples, voting rule, thinking budget) in the .eval_results notes and on the model card, and we publish single-sample scores alongside the maj@k results. That is consistent with the leaderboard rules, which accept self-reported results under a disclosed protocol.

Despite this, the same claims keep being repeated in a way that disrupts the discussion, so we have blocked the account. We'll follow whatever labeling the team decides on.

This discussion seems to be getting unnecessarily confrontational, so I’d like to add a point. In my view, using multiple solutions does not, by itself, make the model card’s reporting approach illegitimate.
Single-response scores and results that aggregate multiple solutions should certainly be distinguished. However, consensus voting and learned answer selection have also been used by major research labs.

  • OpenAI o1 reported AIME results using consensus over 64 solutions, as well as results obtained by reranking 1,000 solutions with a learned scoring function.
  • Anthropic Claude 3.7 Sonnet’s 84.8% GPQA result used compute equivalent to 256 independent attempts and a trained scoring model.
  • xAI Grok 3’s 93.3% on AIME 2025 was a cons@64 result. The GPQA figure on that same page, however, does not specify its evaluation method.
    The key question is therefore not simply whether multiple attempts were used, but whether the evaluation conditions were disclosed transparently. Single-response, majority-vote, and learned-selection results should not be presented as directly comparable under identical conditions. Equally, a disclosed evaluation method should not automatically be treated as misconduct.
    According to the Darwin-180B-RSI submitters, every submitted score documents the sample count, voting method, and thinking budget in both the .eval_results notes and the model card. They also report that the evaluation files for 14 other top GPQA Diamond entries did not specify their evaluation methods. The moderation team should be able to verify that directly.
    If the leaderboard accepts self-reported evaluations, submissions that disclose their methods should be reviewed under the same standards as everyone else. The submitters have also stated that they are willing to follow whatever labeling convention the team requires. I hope the discussion can move toward clear, consistent disclosure standards for all models rather than unnecessary accusations.

@SeaWolf-AI @Sunghokim Replying on the substance.

  1. This report is not about tone. Nothing in it is an insult, and the questions in it don't depend on anything said elsewhere.

  2. Protocol disclosure is not what I'm disputing. The .eval_results notes do state it. My point is that a maj@16 result (GPQA 94.44) and a maj@3 result (MMMU-Pro 79.48) sit in the same ranking as other models' single-sample scores, and the card headlines them as "#1". Sunghokim's comment says single-response and majority-vote results "should not be presented as directly comparable". That is exactly the request in this report.

  3. "We publish single-sample scores alongside the maj@k results." For this model, gpqa_diamond.yaml holds a single value: 94.44, majority vote over up to 16 samples. mmmu_pro.yaml holds only the maj@3 value. Only MMLU-Pro is single-sample. If single-sample GPQA and MMMU-Pro numbers for Darwin-180B-RSI exist, please link them.

  4. Still unanswered from the original thread:

    • What does Qwen3.8-Flash-Next score on GPQA Diamond under the same maj@16, 131K-thinking protocol?
    • What is Darwin-180B-RSI's GPQA Diamond score single-sample, with seed count and confidence interval?
    • Where are the eval harness and per-question outputs for this model?
      The only same-protocol comparison on the card (MMLU-Pro, 88.04 vs 88.12) is +0.08, about 10 questions out of 12,032.
  5. On the labs cited: those numbers were published with the method stated next to them and, as far as I know, alongside single-sample results, not as a single ranking mixing both. And if 14 other top GPQA entries don't specify their method, that is an argument for standardized protocol labels on the leaderboard, not for leaving this entry as an unlabeled #1.

Since the publisher says they'll follow whatever labeling the team decides, I assume there's no objection to separating or labeling entries by protocol (single-sample vs. maj@k). That's all I'm asking the HF team to review.

To hugging face team:

dedeprogames has been blocked for life for abuse at my repo [398 models], on discord and other discussion boards.
I don't take this action lightly, and in 3 years - and close to 100 million downloads in that time - I can count on one hand how many people that have been blocked.

FINAL_Bench org

To hugging face team:

dedeprogames has been blocked for life for abuse at my repo [398 models], on discord and other discussion boards.
I don't take this action lightly, and in 3 years - and close to 100 million downloads in that time - I can count on one hand how many people that have been blocked.

Thank you, @DavidAU . We appreciate you sharing that. We'll leave it with the Hugging Face team.

Thank you @DavidAU for weighing in. Reading your note, it seems this isn't the first time the same account has caused trouble across the community. Open discussion only works when criticism stays on the work and off the people, and I think most readers here would agree with that.

To the Hugging Face team (cc @DavidAU @SeaWolf-AI @Sunghokim ):

  1. Scope. This report concerns one thing: a maj@16 self-reported score ranked #1 next to single-sample scores. A block on another platform, or a dispute about someone's conduct elsewhere, does not answer that question. @DavidAU , your comment is about a separate dispute and doesn't address the leaderboard issue. Please keep to the technical point or leave it to the team.

  2. Same gap between what was measured and what is claimed, documented on other cards. I published a technical audit of DavidAU and Nightmedia model cards: https://huggingface.co/blog/DedeProGames/when-benchmark-numbers-become-marketing

It does not claim any number is fabricated. It shows real numbers being turned into conclusions they don't support:

  • A 7-task legacy suite (ARC, BoolQ, HellaSwag, OBQA, PIQA, WinoGrande) presented as broad intelligence; ARC-C β‰ˆ 0.700 tied to an "OpenAI/Claude/Gemini zone" with no such standard.
  • "Fully uncensored" next to 6/100 refusals on the same card.
  • "Safe at 2M" context backed only by short-context tasks and a 512-token perplexity run. The default mlx_lm.perplexity uses the Tulu-3 SFT train split, concatenates unrelated examples, reports token-level standard errors, and its tokens/sec is evaluation throughput, not generation speed.
  • "Standard MLX metrics": mlx_lm.evaluate defines no fixed suite (--tasks is chosen by the evaluator); columns labeled acc_norm aren't all acc_norm; lm-eval versions in filenames don't match "latest".
  • Universal "+2-4%" for imatrix and "thinking > instruct" with no paired tests.
  • "Mathematical saturation ceiling", "implicit inner state", and LLM-generated "reviews" used as evidence.
  • Agent-positioned models with no SWE-bench, Terminal-Bench or tool-use evals, while the base model's card reports them.
    The article also lists what a convincing evaluation would publish.
  1. Darwin-180B-RSI shows the same gap: maj@16 GPQA (94.44) headlined as "#1"; no parent under the same protocol; no single-sample GPQA for this model; no harness or per-question outputs; +0.08 on MMLU-Pro, the only same-protocol comparison.

  2. Timeline, for the team. I asked technical questions in discussion #1; it was closed without answering them. I filed this report; the publisher then announced a block on my account. Earlier, after I posted methodological criticism on a DavidAU repo, I lost access to it and got a 403 ("not allowed to interact with that user"), documented in the article. I can't prove motive, but each block followed criticism of the claims, and it removes my ability to answer in threads where I'm being described. Some of my comments in #1 were harsh; this report and the article contain no insults.

I ask the team to review the technical questions on their merits, and the timeline as documented.

@DedeProGames i think its fine honestly i dont rlly care about the benchmark scores, i care about the actual practical use and if this model is super useful irl then im fine with it

To clarify the scope: I'm not disputing the benchmark content here. My point is how the results are presented. A maj@16 score is entered in the official tables next to single-sample results and headlined as "#1".

If that's acceptable, the ranking stops measuring the model and starts measuring inference budget: any model can be resubmitted as maj@64 or best-of-N and move up the table. The effect is large. The publisher confirmed 72.85 single-sample for Darwin-27B-Opus, while the card's family table lists 86.9.

The fix is simple: label or separate entries by protocol (single-sample vs. maj@k). The publisher and @Sunghokim both said that's acceptable. That is all this report asks the HF team to review.

It is clear "dedeprogames" is continuing to use an AI(s) (as he does not talk like this) to cause undue issues to the community.
Likewise it appears the AI(s) he is using is hallucinating just as much as he is.

It is also clear, that marking this as "spam" is a desperate attempt to undermine the community, bully @SeaWolf-AI , misdirect the discussion and waste huggingface's abuse teams time.
Gaslighting and bullying at its finest.

@DavidAU Three points.

  1. "The AI is hallucinating": then name one factual error. Not tone, not tooling: one claim in this report or in the article that is wrong. How I draft says nothing about whether the numbers on these cards are what the article says they are.

  2. "Spam": the title "🚩 Report: Spam" is also on #10, opened by a different user. It looks like the default title of HF's report form, not a statement about anyone. The content is a protocol question: maj@16 ranked next to single-sample scores. That isn't bullying.

  3. Since you chose to speak here, this is what the audit documents about your models (https://huggingface.co/blog/DedeProGames/when-benchmark-numbers-become-marketing). It does not say any number is fabricated:

  • Fable Fusion 711: ARC-C ~0.700 tied to an "OpenAI, Claude and Gemini zone of intelligence" (no such threshold exists); "strongest/smartest" claims backed by 7 legacy multiple-choice tasks; no regression tests against the base model's SWE-bench, Terminal-Bench, GPQA or LiveCodeBench.
  • NEO imatrix "+2-4% accuracy" and better long context: no paired controlled tests.
  • "Thinking mode will exceed instruct" and "BF16 2-5 points higher": the tables backing them aren't published; estimates stand in for measurements.
  • The Defiant 9B: "Fully uncensored" while the same card reports 6/100 refusals; "extreme intelligence" and "superior instruction following" supported only by the same 7 tasks; ARC-C presented as "intelligence".
  • 40B Grand Intelligence: claims to beat the 27B by a large margin, yet BoolQ is lower (0.908 vs 0.910); "SOTA IQ/power at 40B" with no defined metric; "human testing" cited as evidence with no protocol.
  • Moderation: after I posted methodological criticism on your repo, I lost access to the discussion and got a 403 ("not allowed to interact with that user"); my request for SWE-bench results went unanswered.

Blocking a critic doesn't rebut any of these points. If one is wrong, say which and why. Otherwise, this thread is about one question: should maj@k and single-sample results be ranked together without labels?

I think some one has to benchmark the model itself .. that would be better than fighting here.

well its a 190B model, its not that easy

As longs as it’s useful I don’t give crap about benchmarks

Sign up or log in to comment