PaSaMaster Ranker

PaSaMaster Ranker is the evidence-grounded paper-ranking model used in PaSaMaster, a self-evolving agentic system for scientific literature retrieval. Built on Qwen3-30B-A3B, it is PaSaMaster's verification and ranking component—not a standalone search engine or general-purpose chat model.

How it works

Given a research intent, a query-specific relevance checklist, and a candidate paper's metadata, abstract, and retrieved evidence, the Ranker:

  • scores each checklist criterion from 1 to 5;
  • provides evidence-grounded rationales and flags weak matches;
  • estimates holistic relevance and reranks verified papers.

This design ranks authentic, traceable paper records instead of generating citations. The model was trained through multidisciplinary knowledge distillation on 42,762 query–paper pairs spanning 19 disciplines and 97 fine-grained topics, including positive, partial-match, and hard-negative examples.

PaSaMaster results

The following are end-to-end PaSaMaster system results using this Ranker, measured on PaSaMaster-Bench (244 expert-curated tasks across 38 disciplines):

Method NDCG@20 Recall@20 Precision@20 F1@20 Hallucination Cost/query
Google Scholar 2.07 1.69 1.48 1.39 0% —
OpenScholar 14.61 11.68 8.52 7.92 0% —
Bohrium Science Navigator 22.39 19.37 12.50 12.26 0% —
DeepSeek-v3.2 35.82 24.76 15.35 15.56 12.94% $0.28
Kimi-K2.5 37.80 28.08 16.95 17.36 26.59% $0.16
MiniMax-M2.7 30.70 24.23 14.42 15.11 32.66% $0.18
GLM-5 35.89 28.99 16.93 18.18 21.64% $0.56
Gemini-3.1-pro 31.34 21.30 11.68 12.48 27.54% $0.38
GPT-5.2 31.59 25.32 16.82 16.69 5.65% $6.06
Google Scholar Labs 30.54 29.01 18.79 18.87 0% —
PaSaMaster 39.52 33.24 23.46 23.00 0% $0.05

PaSaMaster achieves 16.5× the F1@20 of Google Scholar and 37.8% higher F1@20 than GPT-5.2 in the reported evaluation, at approximately $0.05 per query.

Intended use

Use this checkpoint within the PaSaMaster pipeline or a compatible evidence-grounded literature-ranking workflow. Supply retrieved paper evidence and treat its outputs as relevance judgments, not as independent proof that a citation is valid.

Resources

Downloads last month
325
Safetensors
Model size
31B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for PaSaMaster/Ranker

Finetuned
(75)
this model

Paper for PaSaMaster/Ranker