baseten-admin commited on
Commit
e8ea46d
Β·
verified Β·
1 Parent(s): 3e27bfc

Upload folder using huggingface_hub

Browse files
.gitattributes CHANGED
@@ -35,3 +35,4 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
36
  sym_smoothquant/qwen_fp4_exclude_none_algo_smoothquant.png filter=lfs diff=lfs merge=lfs -text
37
  sym_max/qwen_fp4_exclude_none_algo_max.png filter=lfs diff=lfs merge=lfs -text
 
 
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
36
  sym_smoothquant/qwen_fp4_exclude_none_algo_smoothquant.png filter=lfs diff=lfs merge=lfs -text
37
  sym_max/qwen_fp4_exclude_none_algo_max.png filter=lfs diff=lfs merge=lfs -text
38
+ asym_max/qwen_fp4_exclude_none_algo_max_asymmetric.png filter=lfs diff=lfs merge=lfs -text
asym_max/inference.log ADDED
@@ -0,0 +1,61 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
0
  0%| | 0/30 [00:00<?, ?it/s]Loading extension modelopt_cuda_ext_fp8...
 
 
 
1
  3%|β–Ž | 1/30 [00:01<00:32, 1.12s/it]
2
  7%|β–‹ | 2/30 [00:01<00:20, 1.38it/s]
3
  10%|β–ˆ | 3/30 [00:02<00:16, 1.67it/s]
4
  13%|β–ˆβ–Ž | 4/30 [00:02<00:13, 1.86it/s]
5
  17%|β–ˆβ–‹ | 5/30 [00:02<00:12, 1.98it/s] Step 6/30
 
6
  20%|β–ˆβ–ˆ | 6/30 [00:03<00:11, 2.06it/s]
7
  23%|β–ˆβ–ˆβ–Ž | 7/30 [00:03<00:10, 2.11it/s]
8
  27%|β–ˆβ–ˆβ–‹ | 8/30 [00:04<00:10, 2.15it/s]
9
  30%|β–ˆβ–ˆβ–ˆ | 9/30 [00:04<00:09, 2.18it/s]
10
  33%|β–ˆβ–ˆβ–ˆβ–Ž | 10/30 [00:05<00:09, 2.19it/s] Step 11/30
 
11
  37%|β–ˆβ–ˆβ–ˆβ–‹ | 11/30 [00:05<00:08, 2.21it/s]
12
  40%|β–ˆβ–ˆβ–ˆβ–ˆ | 12/30 [00:06<00:08, 2.21it/s]
13
  43%|β–ˆβ–ˆβ–ˆβ–ˆβ–Ž | 13/30 [00:06<00:07, 2.22it/s]
14
  47%|β–ˆβ–ˆβ–ˆβ–ˆβ–‹ | 14/30 [00:06<00:07, 2.22it/s]
15
  50%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆ | 15/30 [00:07<00:06, 2.23it/s] Step 16/30
 
16
  53%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–Ž | 16/30 [00:07<00:06, 2.23it/s]
17
  57%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–‹ | 17/30 [00:08<00:05, 2.23it/s]
18
  60%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ | 18/30 [00:08<00:05, 2.23it/s]
19
  63%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–Ž | 19/30 [00:09<00:04, 2.23it/s]
20
  67%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–‹ | 20/30 [00:09<00:04, 2.23it/s] Step 21/30
 
21
  70%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ | 21/30 [00:10<00:04, 2.23it/s]
22
  73%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–Ž | 22/30 [00:10<00:03, 2.23it/s]
23
  77%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–‹ | 23/30 [00:10<00:03, 2.23it/s]
24
  80%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ | 24/30 [00:11<00:02, 2.23it/s]
25
  83%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–Ž | 25/30 [00:11<00:02, 2.23it/s] Step 26/30
 
26
  87%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–‹ | 26/30 [00:12<00:01, 2.23it/s]
27
  90%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ | 27/30 [00:12<00:01, 2.23it/s]
28
  93%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–Ž| 28/30 [00:13<00:00, 2.24it/s]
29
  97%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–‹| 29/30 [00:13<00:00, 2.24it/s] Step 30/30
 
 
 
 
1
+ /workspace/Model-Optimizer/.venv/lib/python3.12/site-packages/torch/cuda/__init__.py:63: FutureWarning: The pynvml package is deprecated. Please install nvidia-ml-py instead. If you did not install pynvml directly, please report this to the maintainers of the package that installed pynvml for you.
2
+ import pynvml # type: ignore[import]
3
+ Multiple distributions found for package modelopt. Picked distribution: nvidia-modelopt
4
+
5
+ ================================================================================
6
+ NVFP4_GEMM KERNEL STATUS
7
+ ================================================================================
8
+ ❌ NOT AVAILABLE
9
+ Reason: TensorRT-LLM not installed: libmpi.so.40: cannot open shared object file: No such file or directory
10
+
11
+ If you use --restore-from with a COMPRESSED checkpoint,
12
+ this script will EXIT with an error.
13
+
14
+ Use FAKE quantized checkpoints (without --compress) instead
15
+ for quality testing at BF16 speed.
16
+ ================================================================================
17
+
18
+ [22:35:44] GPU: NVIDIA B200 (178.4 GB)
19
+ [22:35:44] Loading Qwen-Image-2512 pipeline from HuggingFace...
20
+
21
+
22
+
23
+
24
+
25
+
26
+ [22:35:45] Pipeline components:
27
+ [22:35:45] - vae: AutoencoderKLQwenImage (126,892,531 params)
28
+ [22:35:45] - text_encoder: Qwen2_5_VLForConditionalGeneration (8,292,166,656 params)
29
+ [22:35:45] - tokenizer: Qwen2Tokenizer
30
+ [22:35:45] - transformer: QwenImageTransformer2DModel (20,430,401,088 params)
31
+ [22:35:45] - scheduler: FlowMatchEulerDiscreteScheduler
32
+ [22:35:45] Backbone model: QwenImageTransformer2DModel
33
+ [22:35:45] Restoring quantized checkpoint from ./experiments_v3_asymmetric/exclude_none_algo_max_asymmetric/qwen_fp4_exclude_none_algo_max_asymmetric.pt...
34
+ Inserted 2838 quantizers
35
+ [22:36:03] Checkpoint restored in 18.6s
36
+ [22:36:03] Moving pipeline to CUDA...
37
+ [22:36:09] Pipeline ready!
38
+ [22:36:09]
39
+ [22:36:09] ============================================================
40
+ [22:36:09] ℹ️ Running with FAKE QUANTIZATION
41
+ [22:36:09] Quality = same as real FP4
42
+ [22:36:09] Speed = BF16 (no speedup)
43
+ [22:36:09] This is fine for quality evaluation.
44
+ [22:36:09] ============================================================
45
+ [22:36:09]
46
+ Generating image with prompt: 'a photo of an astronaut riding a horse on mars, highly detailed, 4k'
47
+ [22:36:09] Generating image with prompt: 'a photo of an astronaut riding a horse on mars, highly detai...'
48
+ [22:36:09] Using negative prompt: 'δ½Žεˆ†θΎ¨ηŽ‡οΌŒδ½Žη”»θ΄¨οΌŒθ‚’δ½“η•Έε½’οΌŒζ‰‹ζŒ‡η•Έε½’οΌŒη”»ι’θΏ‡ι₯±ε’ŒοΌŒθœ‘εƒζ„ŸοΌŒδΊΊθ„Έζ— η»†θŠ‚οΌŒθΏ‡εΊ¦ε…‰ζ»‘οΌŒ...'
49
+ guidance_scale is passed as 4.0, but ignored since the model is not guidance-distilled.
50
+
51
  0%| | 0/30 [00:00<?, ?it/s]Loading extension modelopt_cuda_ext_fp8...
52
+ Loaded extension modelopt_cuda_ext_fp8 in 0.0 seconds
53
+ Step 1/30
54
+
55
  3%|β–Ž | 1/30 [00:01<00:32, 1.12s/it]
56
  7%|β–‹ | 2/30 [00:01<00:20, 1.38it/s]
57
  10%|β–ˆ | 3/30 [00:02<00:16, 1.67it/s]
58
  13%|β–ˆβ–Ž | 4/30 [00:02<00:13, 1.86it/s]
59
  17%|β–ˆβ–‹ | 5/30 [00:02<00:12, 1.98it/s] Step 6/30
60
+
61
  20%|β–ˆβ–ˆ | 6/30 [00:03<00:11, 2.06it/s]
62
  23%|β–ˆβ–ˆβ–Ž | 7/30 [00:03<00:10, 2.11it/s]
63
  27%|β–ˆβ–ˆβ–‹ | 8/30 [00:04<00:10, 2.15it/s]
64
  30%|β–ˆβ–ˆβ–ˆ | 9/30 [00:04<00:09, 2.18it/s]
65
  33%|β–ˆβ–ˆβ–ˆβ–Ž | 10/30 [00:05<00:09, 2.19it/s] Step 11/30
66
+
67
  37%|β–ˆβ–ˆβ–ˆβ–‹ | 11/30 [00:05<00:08, 2.21it/s]
68
  40%|β–ˆβ–ˆβ–ˆβ–ˆ | 12/30 [00:06<00:08, 2.21it/s]
69
  43%|β–ˆβ–ˆβ–ˆβ–ˆβ–Ž | 13/30 [00:06<00:07, 2.22it/s]
70
  47%|β–ˆβ–ˆβ–ˆβ–ˆβ–‹ | 14/30 [00:06<00:07, 2.22it/s]
71
  50%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆ | 15/30 [00:07<00:06, 2.23it/s] Step 16/30
72
+
73
  53%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–Ž | 16/30 [00:07<00:06, 2.23it/s]
74
  57%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–‹ | 17/30 [00:08<00:05, 2.23it/s]
75
  60%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ | 18/30 [00:08<00:05, 2.23it/s]
76
  63%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–Ž | 19/30 [00:09<00:04, 2.23it/s]
77
  67%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–‹ | 20/30 [00:09<00:04, 2.23it/s] Step 21/30
78
+
79
  70%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ | 21/30 [00:10<00:04, 2.23it/s]
80
  73%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–Ž | 22/30 [00:10<00:03, 2.23it/s]
81
  77%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–‹ | 23/30 [00:10<00:03, 2.23it/s]
82
  80%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ | 24/30 [00:11<00:02, 2.23it/s]
83
  83%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–Ž | 25/30 [00:11<00:02, 2.23it/s] Step 26/30
84
+
85
  87%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–‹ | 26/30 [00:12<00:01, 2.23it/s]
86
  90%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ | 27/30 [00:12<00:01, 2.23it/s]
87
  93%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–Ž| 28/30 [00:13<00:00, 2.24it/s]
88
  97%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–‹| 29/30 [00:13<00:00, 2.24it/s] Step 30/30
89
+
90
+ [22:36:23] Image generated in 14.8s, saved as ./experiments_v3_asymmetric/exclude_none_algo_max_asymmetric/qwen_fp4_exclude_none_algo_max_asymmetric.png
91
+ [22:36:23] Done!
asym_max/quantize.log ADDED
The diff for this file is too large to render. See raw diff
 
asym_max/qwen_fp4_exclude_none_algo_max_asymmetric.png ADDED

Git LFS Details

  • SHA256: 70264738dc0aaaff6be59104b399f602c948387165ccf709c95e6982d5ef1d6e
  • Pointer size: 132 Bytes
  • Size of remote file: 1.41 MB
asym_max/qwen_fp4_exclude_none_algo_max_asymmetric.pt ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:215300c35e5a0f2b7097896bce37f7802854f4b3ba25ef38e53a50e49d9303a2
3
+ size 40875159450