=== 2026-09-10T08:59:18-07:00 START v2-31B-heretic (post config patch) [notice] To update, run: python3.12 -m pip install --upgrade pip loading /tank/aimodels/G4-MeroMero-v2-31B-heretic-bf16 Loading weights: 100%|██████████| 1188/1188 [00:00<00:00, 2518.29it/s] building calibration (<= 512 @ seq 8192) 512 calibration rows NVFP4 oneshot: scheme=NVFP4A16, Linear-only, vision/audio/projector/embed/lm_head/norms kept BF16 2026-09-10T15:59:48.3371 | __init__ | WARNING - Disabling tokenizer parallelism due to threading conflict between FastTokenizer and Datasets. Set TOKENIZERS_PARALLELISM=false to suppress this warning. Traceback (most recent call last): File "/usr/local/lib/python3.12/dist-packages/llmcompressor/entrypoints/utils.py", line 68, in pre_process model_args.processor = initialize_processor_from_path( ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "/usr/local/lib/python3.12/dist-packages/llmcompressor/entrypoints/utils.py", line 184, in initialize_processor_from_path processor = AutoProcessor.from_pretrained( ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "/usr/local/lib/python3.12/dist-packages/transformers/models/auto/processing_auto.py", line 327, in from_pretrained return processor_class.from_pretrained( ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "/usr/local/lib/python3.12/dist-packages/transformers/processing_utils.py", line 1715, in from_pretrained args = cls._get_arguments_from_pretrained(pretrained_model_name_or_path, processor_dict, **kwargs) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "/usr/local/lib/python3.12/dist-packages/transformers/processing_utils.py", line 1844, in _get_arguments_from_pretrained sub_processor = auto_processor_class.from_pretrained( ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "/usr/local/lib/python3.12/dist-packages/transformers/models/auto/feature_extraction_auto.py", line 300, in from_pretrained config_dict, _ = FeatureExtractionMixin.get_feature_extractor_dict(pretrained_model_name_or_path, **kwargs) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "/usr/local/lib/python3.12/dist-packages/transformers/feature_extraction_utils.py", line 529, in get_feature_extractor_dict raise OSError( OSError: Can't load feature extractor for '/tank/aimodels/G4-MeroMero-v2-31B-heretic-bf16'. If you were trying to load it from 'https://huggingface.co/models', make sure you don't have a local directory with the same name. Otherwise, make sure '/tank/aimodels/G4-MeroMero-v2-31B-heretic-bf16' is the correct path to a directory containing a preprocessor_config.json file The above exception was the direct cause of the following exception: Traceback (most recent call last): File "/tank/aimodels/meromero-v2-nvfp4-work/quant_nvfp4_gemma.py", line 128, in sys.exit(main()) ^^^^^^ File "/tank/aimodels/meromero-v2-nvfp4-work/quant_nvfp4_gemma.py", line 111, in main oneshot( File "/usr/local/lib/python3.12/dist-packages/llmcompressor/entrypoints/oneshot.py", line 468, in oneshot one_shot = Oneshot(**local_args, **kwargs) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "/usr/local/lib/python3.12/dist-packages/llmcompressor/entrypoints/oneshot.py", line 178, in __init__ pre_process(model_args, dataset_args, output_dir) File "/usr/local/lib/python3.12/dist-packages/llmcompressor/entrypoints/utils.py", line 73, in pre_process raise RuntimeError( RuntimeError: An error occurred when attempting to initialize model processor, which is required when a dataset is provided. To resolve, create and pass in a processor directly to `oneshot`/`train`. === 2026-09-10T08:59:54-07:00 END rc=0 size=512 === 2026-09-10T09:21:40-07:00 START v2-31B-heretic (config patched + processor + data-free) [notice] To update, run: python3.12 -m pip install --upgrade pip usage: ref [-h] --model MODEL --out OUT [--scheme SCHEME] ref: error: the following arguments are required: --model, --out === 2026-09-10T09:21:56-07:00 END rc=2 size=512 === 2026-09-10T09:23:50-07:00 START v2-31B-heretic (config patched + processor + data-free) Unable to find image 'vllm/vllm-openai:latest' locally latest: Pulling from vllm/vllm-openai cf57d2112d89: Already exists c567a87f21d2: Already exists 0b3b5bd92824: Already exists 70f87b6ed43e: Already exists 526d5438c009: Already exists 79ed78d42ca9: Pulling fs layer 5785dfb2d94d: Pulling fs layer 8136a02ba8b6: Pulling fs layer 1dae32d336bd: Pulling fs layer 4f4fb700ef54: Pulling fs layer 0e1d24786a23: Pulling fs layer c83c61a504db: Pulling fs layer 81801e5f6a47: Pulling fs layer e832d0ac2449: Pulling fs layer f41db59aec9f: Pulling fs layer 4a1facbdf857: Pulling fs layer b1cc0c6d03ea: Pulling fs layer 334119c098d4: Pulling fs layer b14dc82c93d9: Pulling fs layer f599001d1dac: Pulling fs layer 66a08f34da9c: Pulling fs layer 2411167b6874: Pulling fs layer a9b8ef092e47: Pulling fs layer 0f36e99efdcd: Pulling fs layer 32d0568ab58d: Pulling fs layer c83c61a504db: Waiting 81801e5f6a47: Waiting 4693bfabf3cd: Pulling fs layer e832d0ac2449: Waiting 7994811847da: Pulling fs layer d80f1ecbeb8c: Pulling fs layer 1dae32d336bd: Waiting 8294aa869476: Pulling fs layer f41db59aec9f: Waiting d510763bc7fa: Pulling fs layer 4f4fb700ef54: Waiting 2b5be6c4f7e6: Pulling fs layer 0e1d24786a23: Waiting c8097f43c701: Pulling fs layer 4a1facbdf857: Waiting 3c84f5c1ca0b: Pulling fs layer b1cc0c6d03ea: Waiting b989116cb4ee: Pulling fs layer f599001d1dac: Waiting 334119c098d4: Waiting cfe38ec4fb3a: Pulling fs layer 2411167b6874: Waiting b14dc82c93d9: Waiting a9b8ef092e47: Waiting 66a08f34da9c: Waiting 8294aa869476: Waiting d510763bc7fa: Waiting 3c84f5c1ca0b: Waiting 2b5be6c4f7e6: Waiting 4693bfabf3cd: Waiting 0f36e99efdcd: Waiting c8097f43c701: Waiting b989116cb4ee: Waiting 32d0568ab58d: Waiting cfe38ec4fb3a: Waiting 7994811847da: Waiting d80f1ecbeb8c: Waiting 79ed78d42ca9: Verifying Checksum 79ed78d42ca9: Download complete 5785dfb2d94d: Verifying Checksum 5785dfb2d94d: Download complete 79ed78d42ca9: Pull complete 5785dfb2d94d: Pull complete 4f4fb700ef54: Verifying Checksum 4f4fb700ef54: Download complete 0e1d24786a23: Verifying Checksum 0e1d24786a23: Download complete c83c61a504db: Verifying Checksum c83c61a504db: Download complete 81801e5f6a47: Verifying Checksum 81801e5f6a47: Download complete e832d0ac2449: Verifying Checksum e832d0ac2449: Download complete f41db59aec9f: Download complete 4a1facbdf857: Verifying Checksum 4a1facbdf857: Download complete b1cc0c6d03ea: Verifying Checksum b1cc0c6d03ea: Download complete 8136a02ba8b6: Download complete 8136a02ba8b6: Pull complete b14dc82c93d9: Verifying Checksum b14dc82c93d9: Download complete f599001d1dac: Download complete 66a08f34da9c: Verifying Checksum 66a08f34da9c: Download complete 2411167b6874: Verifying Checksum 2411167b6874: Download complete a9b8ef092e47: Verifying Checksum a9b8ef092e47: Download complete 0f36e99efdcd: Verifying Checksum 0f36e99efdcd: Download complete 32d0568ab58d: Verifying Checksum 32d0568ab58d: Download complete 4693bfabf3cd: Verifying Checksum 4693bfabf3cd: Download complete 7994811847da: Verifying Checksum 7994811847da: Download complete d80f1ecbeb8c: Verifying Checksum d80f1ecbeb8c: Download complete 8294aa869476: Verifying Checksum 8294aa869476: Download complete 1dae32d336bd: Verifying Checksum 1dae32d336bd: Download complete d510763bc7fa: Verifying Checksum d510763bc7fa: Download complete c8097f43c701: Verifying Checksum c8097f43c701: Download complete 3c84f5c1ca0b: Verifying Checksum 3c84f5c1ca0b: Download complete b989116cb4ee: Verifying Checksum b989116cb4ee: Download complete cfe38ec4fb3a: Download complete 2b5be6c4f7e6: Verifying Checksum 2b5be6c4f7e6: Download complete 1dae32d336bd: Pull complete 4f4fb700ef54: Pull complete 0e1d24786a23: Pull complete c83c61a504db: Pull complete 81801e5f6a47: Pull complete e832d0ac2449: Pull complete f41db59aec9f: Pull complete 4a1facbdf857: Pull complete b1cc0c6d03ea: Pull complete 334119c098d4: Verifying Checksum 334119c098d4: Download complete 334119c098d4: Pull complete b14dc82c93d9: Pull complete f599001d1dac: Pull complete 66a08f34da9c: Pull complete 2411167b6874: Pull complete a9b8ef092e47: Pull complete 0f36e99efdcd: Pull complete 32d0568ab58d: Pull complete 4693bfabf3cd: Pull complete 7994811847da: Pull complete d80f1ecbeb8c: Pull complete 8294aa869476: Pull complete d510763bc7fa: Pull complete 2b5be6c4f7e6: Pull complete c8097f43c701: Pull complete 3c84f5c1ca0b: Pull complete b989116cb4ee: Pull complete cfe38ec4fb3a: Pull complete Digest: sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1 Status: Downloaded newer image for vllm/vllm-openai:latest WARNING: Running pip as the 'root' user can result in broken permissions and conflicting behaviour with the system package manager, possibly rendering your system unusable. It is recommended to use a virtual environment instead: https://pip.pypa.io/warnings/venv. Use the --root-user-action option if you know what you are doing and want to suppress this warning. loading /tank/aimodels/G4-MeroMero-v2-31B-heretic-bf16 Traceback (most recent call last): File "/q/quant_a16_datafree.py", line 36, in model, tok = ref.load_model(a.model) ^^^^^^^^^^^^^^^^^^^^^^^ File "/tank/aimodels/meromero-v2-nvfp4-work/quant_nvfp4_gemma.py", line 73, in load_model tok = AutoTokenizer.from_pretrained(path, trust_remote_code=True) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "/usr/local/lib/python3.12/dist-packages/transformers/models/auto/tokenization_auto.py", line 747, in from_pretrained config = AutoConfig.from_pretrained( ^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "/usr/local/lib/python3.12/dist-packages/transformers/models/auto/configuration_auto.py", line 419, in from_pretrained return config_class.from_dict(config_dict, **unused_kwargs) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "/usr/local/lib/python3.12/dist-packages/transformers/configuration_utils.py", line 878, in from_dict config = cls(**config_dict) ^^^^^^^^^^^^^^^^^^ File "/usr/local/lib/python3.12/dist-packages/huggingface_hub/dataclasses.py", line 275, in init_with_validate initial_init(self, *args, **kwargs) # type: ignore [call-arg] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "/usr/local/lib/python3.12/dist-packages/transformers/configuration_utils.py", line 131, in __init__ self.__post_init__(**additional_kwargs) File "/usr/local/lib/python3.12/dist-packages/transformers/models/gemma4/configuration_gemma4.py", line 348, in __post_init__ self.text_config = Gemma4TextConfig(**self.text_config) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "/usr/local/lib/python3.12/dist-packages/huggingface_hub/dataclasses.py", line 276, in init_with_validate cls.validate(self) # type: ignore [attr-defined] ^^^^^^^^^^^^^^^^^^ File "/usr/local/lib/python3.12/dist-packages/huggingface_hub/dataclasses.py", line 251, in validate validator(self) File "/usr/local/lib/python3.12/dist-packages/transformers/configuration_utils.py", line 476, in validate_architecture hasattr(self, "head_dim") File "/usr/local/lib/python3.12/dist-packages/transformers/configuration_utils.py", line 464, in __getattribute__ return super().__getattribute__(key) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "/usr/local/lib/python3.12/dist-packages/transformers/integrations/heterogeneity/configuration_utils.py", line 266, in __getattribute__ raise AmbiguousGlobalPerLayerAttributeError( transformers.integrations.heterogeneity.configuration_utils.AmbiguousGlobalPerLayerAttributeError: 'head_dim' is a per-layer attribute and may vary across layers. Access it via the individual layer configs instead (e.g. config.per_layer_config[i].head_dim). To read the global config value from config.head_dim anyway, set `allow_global_per_layer_attribute_access` to `True` on the config. Warning: only do this if the caller can safely handle heterogeneous configs; code that assumes a homogeneous model may use the global value incorrectly. === 2026-09-10T09:28:02-07:00 END rc=1 size=512 === 2026-09-10T10:40:09-07:00 START v2-31B-heretic attempt 5 (per_layer_config dropped, image pinned) WARNING: Running pip as the 'root' user can result in broken permissions and conflicting behaviour with the system package manager, possibly rendering your system unusable. It is recommended to use a virtual environment instead: https://pip.pypa.io/warnings/venv. Use the --root-user-action option if you know what you are doing and want to suppress this warning. transformers 5.14.1 loading /tank/aimodels/G4-MeroMero-v2-31B-heretic-bf16 Loading weights: 100%|██████████| 1188/1188 [00:00<00:00, 5430.71it/s] NVFP4 oneshot (DATA-FREE): scheme=NVFP4A16, Linear-only, vision/audio/projector/embed/lm_head/norms kept BF16 2026-09-10T17:40:48.5554 | __init__ | WARNING - Disabling tokenizer parallelism due to threading conflict between FastTokenizer and Datasets. Set TOKENIZERS_PARALLELISM=false to suppress this warning. 2026-09-10T17:40:50.1494 | reset | INFO - Compression lifecycle reset 2026-09-10T17:40:50.1782 | from_modifiers | INFO - Creating recipe from modifiers Applying quantization config: 100%|██████████| 410/410 [00:00<00:00, 2863.72it/s] 2026-09-10T17:40:50.3684 | initialize | INFO - Compression lifecycle initialized for 1 modifiers 2026-09-10T17:40:50.3685 | IndependentPipeline | INFO - Inferred `DataFreePipeline` for `QuantizationModifier` 2026-09-10T17:42:45.7622 | finalize | INFO - Compression lifecycle finalized for 1 modifiers saving -> /tank/aimodels/G4-MeroMero-v2-31B-heretic-NVFP4A16 Compressing model: 100%|██████████| 410/410 [00:11<00:00, 36.91it/s] Writing model shards: 100%|██████████| 2/2 [00:08<00:00, 4.01s/it] Dispatching model: 100%|██████████| 1763/1763 [00:00<00:00, 58662.76it/s] DONE === 2026-09-10T10:43:16-07:00 END rc=0 size=19G