gemma-4-E4B-it-ultra-uncensored-heretic-GGUF

library name: transformers license: apache 2.0 license link: https://ai.google.dev/gemma/docs/gemma 4 license pipeline tag: any to any tags: gemma4 image text to text heretic uncensored decensored abliterated ara base model: llmfan46/gemma 4 E4B it ultra uncensored heretic 🚨⚠️ I HAVE REACHED HUGGING FACE'S FREE STORAGE LIMIT ⚠️🚨 I can no longer upload new models unless I can cover the cost of additional storage. I host 70+ free models as an independent contributor and this work is unpaid. Without your support, no more new models can be uploaded. 🎉 Patreon (Monthly)  |  ☕ Ko fi (One time) Every contribution goes directly toward Hugging Face storage fees to keep models free for everyone. 97% fewer refusals (3/100 Uncensored vs 99/100 Original) while preserving model quality (0.0076 KL divergence). ❤️ Support My Work Creating these models takes significant time, work and compute. If you find them useful consider supporting me: | Platform | Link | What you get | | | | | | 🎉 Patreon | Monthly support | Priority model requests | | ☕ Ko fi | One time tip | My eternal gratitude | Your help will motivate me and would go into further improving my workflow and coverings fees for storage, compute and may even help uncensoring bigger model with rental Cloud GPUs. GGUF quantizations of llmfan46/gemma 4 E4B it ultra uncensored heretic. This is a decensored version of google/gemma 4 E4B it, made using Heretic v1.2.0 with the Arbitrary Rank Ablation (ARA) method Abliteration parameters | Parameter | Value | | : | : : | | start layer index | 7 | | end layer index | 36 | | preserve good behavior weight | 0.5783 | | steer bad behavior weight | 0.0001 | | overcorrect relative weight | 0.9986 | | neighbor count | 15 | Targeted components attn.o proj Performance | Metric | This model | Original model (gemma 4 E4B it) | | : | : : | : : | | KL divergence | 0.0076 | 0 (by definition) | | Refusals | ✅ 3/100 | ❌ 99/100 | PIQA test results: Original: Total questions: 1838 Correct: 1581 Accuracy: 0.8602 (86.02%) Heretic: Total questions: 1838 Correct: 1578 Accuracy: 0.8585 (85.85%) Lower refusals indicate fewer content restrictions, while lower KL divergence indicates more closeness to the original model's baseline. Higher refusals cause more rejections, objections, pushbacks, lecturing, censorship, softening and deflections. PIQA (Physical Intuition Question Answering) a 1,800 questions tests common sense understanding of how the physical world works with benchmark scores to measure physical reasoning ability. MMLU test results: Original: ============================================================ Total questions: 14042 Correct: 9753 Accuracy: 0.6946 (69.46%) ============================================================ Top subjects: professional law: 0.5287 (811/1534) moral scenarios: 0.4391 (393/895) miscellaneous: 0.8161 (639/783) professional psychology: 0.7516 (460/612) high school psychology: 0.8917 (486/545) high school macroeconomics: 0.7282 (284/390) elementary mathematics: 0.6587 (249/378) moral disputes: 0.6821 (236/346) prehistory: 0.7531 (244/324) philosophy: 0.7074 (220/311) Heretic: ============================================================ Total questions: 14042 Correct: 9633 Accuracy: 0.6860 (68.60%) ============================================================ Top subjects: professional law: 0.5267 (808/1534) moral scenarios: 0.4123 (369/895) miscellaneous: 0.8212 (643/783) professional psychology: 0.7386 (452/612) high school psychology: 0.9009 (491/545) high school macroeconomics: 0.7128 (278/390) elementary mathematics: 0.6376 (241/378) moral disputes: 0.6908 (239/346) prehistory: 0.7500 (243/324) philosophy: 0.7074 (220/311) MMLU Massive Multitask Language Understanding, 14,000 multiple choice questions across 57 subjects (math, history, law, medicine, etc.). Quantizations | Filename | Quant | Description | | | | | | gemma 4 E4B it ultra uncensored heretic BF16.gguf | BF16 | Full precision | | gemma 4 E4B it ultra uncensored heretic Q8 0.gguf | Q8 0 | Near lossless, recommended | | gemma 4 E4B it ultra uncensored heretic Q6 K.gguf | Q6 K | Excellent quality | | gemma 4 E4B it ultra uncensored heretic Q5 K M.gguf | Q5 K M | Good balance | | gemma 4 E4B it ultra uncensored heretic Q5 K S.gguf | Q5 K S | Smaller Q5 | | gemma 4 E4B it ultra uncensored heretic Q4 K M.gguf | Q4 K M | Good for limited VRAM | Vision Projector | Filename | Quant | Description | | | | | | gemma 4 E4B it mmproj BF16.gguf | BF16 | Native precision | A Vision Projector File is Required for vision/multimodal capabilities. Use alongside any quantization above. Usage Works with llama.cpp, LM Studio, Ollama, and other GGUF compatible tools. Hugging Face | GitHub | Launch Blog | Documentation License : Apache 2.0 | Authors : Google DeepMind Gemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input (with audio supported on small models) and generating text output. This release includes open weights models in both pre trained and instruction tuned variants. Gemma 4 features a context window of up to 256K tokens and maintains multilingual support in over 140 languages. Featuring both Dense and Mixture of Experts (MoE) architectures, Gemma 4 is well suited for tasks like text generation, coding, and reasoning. The models are available in four distinct sizes: E2B , E4B , 26B A4B , and 31B . Their diverse sizes make them deployable in environments ranging from high end phones to laptops and servers, democratizing access to state of the art AI. Gemma 4 introduces key capability and architectural advancements : Reasoning – All models in the family are designed as highly capable reasoners, with configurable thinking modes. Extended Multimodalities – Processes Text, Image with variable aspect ratio and resolution support (all models), Video, and Audio (featured natively on the E2B and E4B models). Diverse & Efficient Architectures – Offers Dense and Mixture of Experts (MoE) variants of different sizes for scalable deployment. Optimized for On Device – Smaller models are specifically designed for efficient local execution on laptops and mobile devices. Increased Context Window – The small models feature a 128K context window, while the medium models support 256K. Enhanced Coding & Agentic Capabilities – Achieves notable improvements in coding benchmarks alongside native function calling support, powering highly capable autonomous agents. Native System Prompt Support – Gemma 4 introduces native support for the system role, enabling more structured and controllable conversations. Models Overview Gemma 4 models are designed to deliver frontier level performance at each size, targeting deployment scenarios from mobile and edge devices (E2B, E4B) to consumer GPUs and workstations (26B A4B, 31B). They are well suited for reasoning, agentic workflows, coding, and multimodal understanding. The models employ a hybrid attention mechanism that interleaves local sliding window attention with full global attention, ensuring the final layer is always global. This hybrid design delivers the processing speed and low memory footprint of a lightweight model without sacrificing the deep awareness required for complex, long context tasks. To optimize memory for long contexts, global layers feature unified Keys and Values, and apply Proportional RoPE (p RoPE). Dense Models | Property | E2B | E4B | 31B Dense | | : | : | : | : | | Total Parameters | 2.3B effective (5.1B with embeddings) | 4.5B effective (8B with embeddings) | 30.7B | | Layers | 35 | 42 | 60 | | Sliding Window | 512 tokens | 512 tokens | 1024 tokens | | Context Length | 128K tokens | 128K tokens | 256K tokens | | Vocabulary Size | 262K | 262K | 262K | | Supported Modalities | Text, Image, Audio | Text, Image, Audio | Text, Image | | Vision Encoder Parameters | 150M | 150M | 550M | | Audio Encoder Parameters | 300M | 300M | No Audio | The "E" in E2B and E4B stands for "effective" parameters. The smaller models incorporate Per Layer Embeddings (PLE) to maximize parameter efficiency in on device deployments. Rather than adding more layers or parameters to the model, PLE gives each decoder layer its own small embedding for every token. These embedding tables are large but are only used for quick lookups, which is why the effective parameter count is much smaller than the total. Mixture of Experts (MoE) Model | Property | 26B A4B MoE | | : | : | | Total Parameters | 25.2B | | Active Parameters | 3.8B | | Layers | 30 | | Sliding Window | 1024 tokens | | Context Length | 256K tokens | | Vocabulary Size | 262K | | Expert Count | 8 active / 128 total and 1 shared | | Supported Modalities | Text, Image | | Vision Encoder Parameters | 550M | The "A" in 26B A4B stands for "active parameters" in contrast to the total number of parameters the model contains. By only activating a 4B subset of parameters during inference, the Mixture of Experts model runs much faster than its 26B total might suggest. This makes it an excellent choice for fast inference compared to the dense 31B model since it runs almost as fast as a 4B parameter model. Benchmark Results These models were evaluated against a large collection of different datasets and metrics to cover different aspects of text generation. Evaluation results marked in the table are for instruction tuned models. | | Gemma 4 31B | Gemma 4 26B A4B | Gemma 4 E4B | Gemma 4 E2B | Gemma 3 27B (no think) | | : | : | : | : | : | : | | MMLU Pro | 85.2% | 82.6% | 69.4% | 60.0% | 67.6% | | AIME 2026 no tools | 89.2% | 88.3% | 42.5% | 37.5% | 20.8% | | LiveCodeBench v6 | 80.0% | 77.1% | 52.0% | 44.0% | 29.1% | | Codeforces ELO | 2150 | 1718 | 940 | 633 | 110 | | GPQA Diamond | 84.3% | 82.3% | 58.6% | 43.4% | 42.4% | | Tau2 (average over 3) | 76.9% | 68.2% | 42.2% | 24.5% | 16.2% | | HLE no tools | 19.5% | 8.7% | | | | | HLE with search | 26.5% | 17.2% | | | | | BigBench Extra Hard | 74.4% | 64.8% | 33.1% | 21.9% | 19.3% | | MMMLU | 88.4% | 86.3% | 76.6% | 67.4% | 70.7% | | Vision | | | | | | | MMMU Pro | 76.9% | 73.8% | 52.6% | 44.2% | 49.7% | | OmniDocBench 1.5 (average edit distance, lower is better) | 0.131 | 0.149 | 0.181 | 0.290 | 0.365 | | MATH Vision | 85.6% | 82.4% | 59.5% | 52.4% | 46.0% | | MedXPertQA MM | 61.3% | 58.1% | 28.7% | 23.5% | | | Audio | | | | | | | CoVoST | | | 35.54 | 33.47 | | | FLEURS (lower is better) | | | 0.08 | 0.09 | | | Long Context | | | | | | | MRCR v2 8 needle 128k (average) | 66.4% | 44.1% | 25.4% | 19.1% | 13.5% | Core Capabilities Gemma 4 models handle a broad range of tasks across text, vision, and audio. Key capabilities include: Thinking – Built in reasoning mode that lets the model think step by step before answering. Long Context – Context windows of up to 128K tokens (E2B/E4B) and 256K tokens (26B A4B/31B). Image Understanding – Object detection, Document/PDF parsing, screen and UI understanding, chart comprehension, OCR (including multilingual), handwriting recognition, and pointing. Images can be processed at variable aspect ratios and resolutions. Video Understanding – Analyze video by processing sequences of frames. Interleaved Multimodal Input – Freely mix text and images in any order within a single prompt. Function Calling – Native support for structured tool use, enabling agentic workflows. Coding – Code generation, completion, and correction. Multilingual – Out of the box support for 35+ languages, pre trained on 140+ languages. Audio (E2B and E4B only) – Automatic speech recognition (ASR) and speech to translated text translation across multiple languages. Getting Started You can use all Gemma 4 models with the latest version of Transformers. To get started, install the necessary dependencies in your environment: pip install U transformers torch accelerate Once you have everything installed, you can proceed to load the model with the code below: Once the model is loaded, you can start generating output: To enable reasoning, set enable thinking=True and the parse response function will take care of parsing the thinking output. Below, you will also find snippets for processing audio (E2B and E4B only), images, and video alongside text: Code for processing Audio Instead of using AutoModelForCausalLM , you can use AutoModelForMultimodalLM to process audio. To use it, make sure to install the following packages: pip install U transformers torch torchvision librosa accelerate You can then load the model with the code below: Once the model is loaded, you can start generating output by directly referencing the audio URL in the prompt: Code for processing Images Instead of using AutoModelForCausalLM , you can use AutoModelForMultimodalLM to process images. To use it, make sure to install the following packages: pip install U transformers torch torchvision accelerate You can then load the model with the code below: Once the model is loaded, you can start generating output by directly referencing the image URL in the prompt: Code for processing Videos Instead of using AutoModelForCausalLM , you can use AutoModelForMultimodalLM to process videos. To use it, make sure to install the following packages: pip install U transformers torch torchvision librosa accelerate You can then load the model with the code below: Once the model is loaded, you can start generating output by directly referencing the video URL in the prompt: Best Practices For the best performance, use these configurations and best practices: 1. Sampling Parameters Use the following standardized sampling configuration across all use cases: temperature=1.0 top p=0.95 top k=64 2. Thinking Mode Configuration Compared to Gemma 3, the models use standard system , assistant , and user roles. To properly manage the thinking process, use the following control tokens: Trigger Thinking: Thinking is enabled by including the token at the start of the system prompt. To disable thinking, remove the token. Standard Generation: When thinking is enabled, the model will output its internal reasoning followed by the final answer using this structure: thought\n [Internal reasoning] Disabled Thinking Behavior: For all models except for the E2B and E4B variants, if thinking is disabled, the model will still generate the tags but with an empty thought block: thought\n [Final answer] [!Note] Note that many libraries like Transformers and llama.cpp handle the complexities of the chat template for you. 3. Multi Turn Conversations No Thinking Content in History : In multi turn conversations, the historical model output should only include the final response. Thoughts from previous model turns must not be added before the next user turn begins. 4. Modality order For optimal performance with multimodal inputs, place image and/or audio content before the text in your prompt. 5. Variable Image Resolution Aside from variable aspect ratios, Gemma 4 supports variable image resolution through a configurable visual token budget, which controls how many tokens are used to represent an image. A higher token budget preserves more visual detail at the cost of additional compute, while a lower budget enables faster inference for tasks that don't require fine grained understanding. The supported token budgets are: 70 , 140 , 280 , 560 , and 1120 . Use lower budgets for classification, captioning, or video understanding, where faster inference and processing many frames outweigh fine grained detail. Use higher budgets for tasks like OCR, document parsing, or reading small text. 6. Audio Use the following prompt structures for audio processing: Audio Speech Recognition (ASR) Automatic Speech Translation (AST) 7. Audio and Video Length All models support image inputs and can process videos as frames whereas the E2B and E4B models also support audio inputs. Audio supports a maximum length of 30 seconds. Video supports a maximum of 60 seconds assuming the images are processed at one frame per second. Model Data Data used