deberta-v3-base-fever

microsoft/deberta-v3-base fine-tuned on pietrolesci/nli_fever for 3-way fact verification, framed as sequence-pair classification (evidence, claim). Full fine-tuning, no adapters or quantization.

Labels

id label
0 SUPPORTS
1 REFUTES
2 NOT ENOUGH INFO

Input format

  • Segment A: evidence (Wikipedia page title + evidence sentences joined with .; may be empty).
  • Segment B: claim.
  • tokenizer(evidence, claim, truncation="only_first", max_length=256) — the claim is never truncated; the tail of the evidence is truncated. Rows whose claim leaves < 32 evidence tokens use longest_first.

Usage

import torch
from transformers import AutoModelForSequenceClassification, AutoTokenizer

repo_id = "thealper2/deberta-v3-base-fever"
tokenizer = AutoTokenizer.from_pretrained(repo_id)
model = AutoModelForSequenceClassification.from_pretrained(repo_id).eval()

evidence = "Paris . Paris is the capital and most populous city of France ."
claim = "Paris is the capital of France."
inputs = tokenizer(evidence, claim, truncation="only_first", max_length=256, return_tensors="pt")
with torch.inference_mode():
    probabilities = model(**inputs).logits.softmax(dim=-1)[0]
print({model.config.id2label[i]: round(p.item(), 4) for i, p in enumerate(probabilities)})

Training data

Dataset field mapping: premise = claim, hypothesis = evidence, fever_gold_label = label. The raw integer label (entailment/neutral/contradiction) was re-mapped to the ids above; verifiable was not used (it is a function of the label).

split source rows SUPPORTS REFUTES NOT ENOUGH INFO
train Hub train minus validation 197,929 117,601 46,507 33,821
validation Hub train, 5% grouped by claim text 10,417 5,993 2,606 1,818
test Hub dev 19,998 6,666 6,666 6,666
  • The Hub test split is unlabeled; the labeled, class-balanced Hub dev split is used as the held-out test set. It was not used for checkpoint selection or hyperparameter tuning.
  • Validation is used for checkpoint selection (best validation macro F1).
  • No deduplication was applied; duplicate/conflicting rows were kept as released.
  • Train: 12,040 duplicate (claim, evidence) rows, 496 pairs with conflicting labels; all 5,194 empty-evidence rows are NOT ENOUGH INFO.
  • Train/test overlap: 3 shared claims, 2 shared (claim, evidence) pairs, 98 shared non-empty evidence strings.

Training procedure

hyperparameter value
epochs 3
learning rate 2e-05
scheduler linear, warmup ratio 0.1
optimizer adamw_torch_fused (weight decay 0.01)
per-device batch size 16
gradient accumulation 2
effective batch size 32
max grad norm 1.0
max sequence length 256 (dynamic padding)
loss cross-entropy (class weighting: none)
precision bf16
seed 42
optimizer steps 18558
hardware NVIDIA GeForce RTX 5060 Ti
training time 0.90 h
parameters 184,424,451
best checkpoint epoch 2.0, validation macro F1 0.8712

Libraries: torch 2.11.0+cu128, transformers 5.17.0, datasets 4.3.0, accelerate 1.12.0, scikit-learn 1.8.0, numpy 2.1.3

Evaluation (held-out test = Hub dev, n=19,998)

metric value
accuracy 0.7768
macro precision 0.7804
macro recall 0.7768
macro F1 0.7773
weighted F1 0.7773
label precision recall F1 support
SUPPORTS 0.8285 0.8639 0.8459 6,666
REFUTES 0.8307 0.7361 0.7806 6,666
NOT ENOUGH INFO 0.6819 0.7304 0.7053 6,666

Confusion matrix (rows = true, columns = predicted):

SUPPORTS REFUTES NOT ENOUGH INFO
SUPPORTS 5,759 154 753
REFUTES 241 4,907 1,518
NOT ENOUGH INFO 951 846 4,869

Slices:

slice n accuracy macro F1
all 19,998 0.7768 0.7773
non_empty_evidence 18,209 0.7918 0.7860
empty_evidence 1,789 0.6244 0.2650
claim_unseen_in_train 19,995 0.7768 0.7772
evidence_not_truncated 19,941 0.7768 0.7773

confusion_matrix

training_curves

Limitations

  • Evidence is the FEVER retrieval output shipped with NLI-FEVER; performance with other retrievers/evidence formats is not measured.
  • Evidence longer than 256 tokens (with the claim) is truncated from the end.
  • In training data, empty evidence always co-occurs with NOT ENOUGH INFO; the model can learn this shortcut (see the empty_evidence slice).
  • Label noise: some identical (claim, evidence) pairs carry different labels in the training data.
  • English Wikipedia claims only; the train (imbalanced) and test (balanced) class priors differ.
Downloads last month
47
Safetensors
Model size
0.2B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for thealper2/deberta-v3-base-fever

Finetuned
(783)
this model

Dataset used to train thealper2/deberta-v3-base-fever

Evaluation results