Atlas3D/JEV-27B-VL-GGUF

GGUF builds of autotrust/JEV-27B-VL for llama.cpp's native decision endpoint /v1/systemone, with JEV's own readout: false/true for yes/no, the digits 0โ€“5 for scores, and letters for choices, plus its decision-head bias and calibrated temperatures. These are also the builds for older GPUs such as V100s, which cannot run FP8/FP4 checkpoints.

Requires a patched llama.cpp (for now). The jev decision type is not in upstream llama.cpp yet. Apply llama.cpp-jev.patch (made against upstream bed0a85) and build. It adds the converter class, the GGUF metadata keys and the server readout.

file size top choice agrees with bf16 (309 decisions) largest prob. change
JEV-27B-VL-BF16.gguf 54 GB 99.4% (307/309): llama.cpp vs vLLM at full precision 0.017
JEV-27B-VL-Q8_0.gguf 29 GB 99.0% (306/309), recommended 0.023
JEV-27B-VL-Q4_K_M.gguf โš ๏ธ 17 GB 92.9% (287/309): less accurate, see below 0.102
mmproj-BF16.gguf, mmproj-Q8_0.gguf 0.9 / 0.6 GB vision projector

โš ๏ธ Q4_K_M is less accurate. About 1 decision in 14 differs from full precision. Prefer Q8_0 when it fits. Q4_K_M is for 24 GB-class cards.

The reference is autotrust's own vLLM server (serve_decide.py) at bf16, on the same recorded requests. Measured on an RTX PRO 6000 Blackwell: about 80 ms per decision for Q8_0, sequential, with prompt caching off. VRAM in use: BF16 56.6 GB, Q8_0 33.6 GB, Q4_K_M 22.7 GB (4k context). V100 (sm_70) builds compile with CUDA 12.x (CUDA 13 dropped sm_70). Results on real V100 hardware have not been measured yet. This is a narrow test of the decision head. Free-text generation was not separately benchmarked.

Serve

llama-server --model JEV-27B-VL-Q8_0.gguf --mmproj mmproj-BF16.gguf --n-gpu-layers 999 \
  --ctx-size 4096 --flash-attn on --jinja --host 127.0.0.1 --port 8080
curl -s localhost:8080/v1/systemone -H 'Content-Type: application/json' -d '{"state": "...",
  "questions": {"q": {"type": "choice", "instructions": "...", "criteria": {"a": null, "b": null}}}}'

License

Apache-2.0, as with the original. This is a modified (converted and quantized) version of autotrust/JEV-27B-VL, which is itself based on Qwen/Qwen3.8-27B. LICENSE is included unchanged. The patch modifies llama.cpp (MIT).

Downloads last month
7,447
GGUF
Model size
27B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

4-bit

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for Atlas3D/JEV-27B-VL-GGUF

Base model

Qwen/Qwen3.8-27B
Quantized
(10)
this model

Space using Atlas3D/JEV-27B-VL-GGUF 1