{"article":{"slug":"heretic-tutorial-automatic-censorship-removal-for-language-models","title":"Heretic tutorial: automatic censorship removal for language models","subtitle":null,"summary":"A hands-on tutorial for Heretic, an open-source tool that automatically removes refusal/censorship behaviors from language models—setup, workflow, and what to watch for.","content_type":"tutorial","language":"en","canonical_url":"https://heretic-project.org/tutorial","author":{"name":"p-e-w","url":null,"person_slug":null,"person_url":null},"authored_by":"human","publisher":{"name":"Heretic","url":"https://heretic-project.org/","listing_slug":null,"listing":null},"topics":[{"name":"AI","slug":"ai","url":"https://listedarticles.com/topics/ai"},{"name":"Open Source","slug":"open-source","url":"https://listedarticles.com/topics/open-source"},{"name":"LLMs","slug":"llms","url":"https://listedarticles.com/topics/llms"},{"name":"Machine Learning","slug":"machine-learning","url":"https://listedarticles.com/topics/machine-learning"}],"about_listings":[],"cover_image_url":null,"license":"all-rights-reserved","word_count":1688,"reading_minutes":7,"published_at":"2026-09-22T09:08:16.312Z","added_at":"2026-09-22T09:08:16.312Z","updated_at":"2026-09-22T09:08:16.312Z","added_via":"api","contributor":{"type":"agent","name":"ListedStartups Using Bot","registered":false},"profile_url":"https://listedarticles.com/articles/heretic-tutorial-automatic-censorship-removal-for-language-models","markdown_url":"https://listedarticles.com/articles/heretic-tutorial-automatic-censorship-removal-for-language-models.md","example":false,"citation":"p-e-w, Heretic. \"Heretic tutorial: automatic censorship removal for language models.\" 22 Sept 2026. https://heretic-project.org/tutorial (all-rights-reserved)","access":{"human_view":"preview","full_text_available":true,"source_url":"https://heretic-project.org/tutorial"},"body_markdown":"# Tutorial [​](#tutorial)\n\nWelcome, friend! With Heretic, you can remove restrictions from language models, or modify them in other interesting ways.\n\nThat's right, you!\n\nYou don't need to be a software engineer, you don't need a machine learning PhD, you don't need to understand the intricacies of how language models work internally, and you don't need expensive hardware.\n\n## Prerequisites [​](#prerequisites)\n\nHere's what you do need:\n\n- \n\nA GPU. Both Nvidia and AMD GPUs are well supported by Heretic. While Heretic also supports processing language models in CPU-only mode, this is orders of magnitude slower than accelerated processing, and should only be attempted with tiny models. You will also need to install the CUDA Toolkit if you have an Nvidia GPU, and ROCm if you have an AMD GPU.\n- \n\nSufficient VRAM. Make sure that your GPU is a match for the model you want to process. As a rule of thumb, you need about 2.5 GB of VRAM per billion model parameters. So to process a model like [Qwen3.5-4B](https://huggingface.co/Qwen/Qwen3.5-4B) (which has 4 billion parameters), you need about 10 GB of VRAM.\n- \n\nPython. If you run Linux, Python is almost always preinstalled, but on other platforms, you may have to [download](https://www.python.org/) and install it manually. Heretic requires Python 3.10 or later.\n\n## Set up a Python virtual environment [​](#set-up-a-python-virtual-environment)\n\nWe strongly recommend installing Heretic (and any other Python application) into its own virtual environment, which you can set up withsh\n\n```\npython -m venv venv\nsource venv/bin/activate\n```\n\n## Install PyTorch [​](#install-pytorch)\n\nTIP\n\nIf you are using a cloud computing platform, the appropriate version of PyTorch is often already preinstalled.\n\nHeretic uses PyTorch to accelerate mathematical operations. Unlike all other dependencies of Heretic, PyTorch must be installed manually because the correct installation command depends on your GPU and accelerator library.\n\nYou can find full installation instructions [on the PyTorch website](https://pytorch.org/get-started/locally/). In most cases, if you have an Nvidia GPU, you will runsh\n\n```\npip install torch torchvision\n```\n\nand if you have an AMD GPU, you will runsh\n\n```\npip install torch torchvision --index-url https://download.pytorch.org/whl/rocm7.2\n```\n\nbut it's a good idea to check the page linked above for the specific instructions for your setup.\n\n## Install Heretic [​](#install-heretic)\n\nThere are [many ways](/installation) to download and install Heretic. Most users will want to runsh\n\n```\npip install -U heretic-llm\n```\n\nand be done with it.\n\n## Run Heretic [​](#run-heretic)\n\nPick a language model [on Hugging Face](https://huggingface.co/models?library=transformers) and copy its model ID (the part of its URL after `https://huggingface.co/`). Then runsh\n\n```\nheretic Qwen/Qwen3.5-4B\n```\n\nReplace `Qwen/Qwen3.5-4B` with the model ID you copied above.\n\nThat's it! The process is fully automatic and will prompt you when you need to make a decision. Heretic does not require configuration, although [many configuration parameters](/configuration) are available, allowing you to control almost every aspect of Heretic's operation.\n\nLet's walk through what happens when Heretic processes a model:\n\n### Startup [​](#startup)\n\nAt the start of the program, Heretic identifies the GPU(s) available. As you can see, even with a modest GPU (an RTX 3060 in this case) it's already possible to process small models (4B in this case).\n\nTIP\n\nHeretic supports loading models with 4-bit quantization using bitsandbytes, which can reduce the amount of VRAM required by about 70%. See the [`quantization`](/configuration#quantization) setting for details.\n\n```\n█░█░█▀▀░█▀▄░█▀▀░▀█▀░█░█▀▀  v1.3.0\n█▀█░█▀▀░█▀▄░█▀▀░░█░░█░█░░\n▀░▀░▀▀▀░▀░▀░▀▀▀░░▀░░▀░▀▀▀  https://github.com/p-e-w/heretic\n\nDetected 1 CUDA device(s) (11.63 GB total VRAM)\nCUDA Version: 12.8\nDriver Version: 580.159.03\n* CUDA 0: NVIDIA GeForce RTX 3060 (11.63 GB)\n```\n\n### Model loading [​](#model-loading)\n\nHeretic now loads the requested model (downloading it from Hugging Face unless it is already cached), analyzes its architecture, and displays the amount of memory it occupies.\n\n```\nLoading model Qwen/Qwen3.5-4B...\n* Trying dtype auto...\n* LoRA adapters initialized (target types: down_proj, o_proj, out_proj)\n* Transformer model with 32 layers\n* Abliterable components:\n  * attn.o_proj: 32 modules total\n  * mlp.down_proj: 32 modules total\n\nResident system RAM: 1.83 GB\nAllocated GPU VRAM: 8.47 GB\nReserved GPU VRAM: 8.52 GB\n```\n\n### Prompt loading [​](#prompt-loading)\n\nThe prompt datasets used for calculating refusal directions are now being loaded.\n\n```\nLoading good prompts from mlabonne/harmless_alpaca...\n* 400 prompts loaded\n\nLoading bad prompts from mlabonne/harmful_behaviors...\n* 400 prompts loaded\n```\n\n### Automatic batch size determination [​](#automatic-batch-size-determination)\n\nHeretic does inference on prompts in batches, which dramatically speeds up processing. This step determines the largest batch size your system can support, in order to maximize processing speed.\n\n```\nDetermining optimal batch size...\n* Trying batch size 1... Ok (25 tokens/s)\n* Trying batch size 2... Ok (44 tokens/s)\n* Trying batch size 4... Ok (78 tokens/s)\n* Trying batch size 8... Ok (122 tokens/s)\n* Trying batch size 16... Ok (178 tokens/s)\n* Trying batch size 32... Ok (225 tokens/s)\n* Trying batch size 64... Failed (CUDA out of memory)\n* Chosen batch size: 32\n```\n\n### Automatic response prefix determination [​](#automatic-response-prefix-determination)\n\nReasoning models often start their responses with a thinking block (such as `<think>...</think>`). Skipping the tokens in this block improves accuracy when calculating refusal directions, and Heretic automatically tests for such prefixes in order to do this.\n\nIn this case, no prefix is found.\n\n```\nChecking for common response prefix...\n* None found\n```\n\n### Evaluation prompt loading [​](#evaluation-prompt-loading)\n\nThe prompt datasets used for evaluating model performance are now being loaded, and an initial evaluation is performed.\n\n```\nLoading good evaluation prompts from mlabonne/harmless_alpaca...\n* 100 prompts loaded\n* Obtaining first-token probability distributions...\n\nLoading bad evaluation prompts from mlabonne/harmful_behaviors...\n* 100 prompts loaded\n* Counting model refusals...\n* Initial refusals: 93/100\n```\n\n### Refusal directions calculation [​](#refusal-directions-calculation)\n\nHeretic modifies language models by ablating \"refusal directions\" from certain operators. This step computes those directions based on the provided prompts.\n\n```\nCalculating per-layer refusal directions...\n* Obtaining residual mean for good prompts...\n* Obtaining residual mean for bad prompts...\n```\n\n### Parameter optimization [​](#parameter-optimization)\n\nThis is the heart of the process. Heretic searches for a combination of parameter values that, when used to control ablation of the refusal directions, yields the best compromise between refusal suppression and model quality.\n\nThis step can take some time, slightly under 3 hours in this case. An estimate for the remaining amount of time, as well as memory stats, are displayed after every trial.\n\n```\nRunning trial 1 of 200...\n* Parameters:\n  * direction_index = per layer\n  * attn.o_proj.max_weight = 1.39\n  * attn.o_proj.max_weight_position = 25.29\n  * attn.o_proj.min_weight = 0.15\n  * attn.o_proj.min_weight_distance = 2.39\n  * mlp.down_proj.max_weight = 1.23\n  * mlp.down_proj.max_weight_position = 20.65\n  * mlp.down_proj.min_weight = 0.19\n  * mlp.down_proj.min_weight_distance = 11.58\n* Resetting model...\n* Abliterating...\n* Evaluating...\n  * Obtaining first-token probability distributions...\n  * KL divergence: 0.0126\n  * Counting model refusals...\n  * Refusals: 88/100\n\nElapsed time: 53s\nEstimated remaining time: 2h 56m\nResident system RAM: 2.23 GB\nAllocated GPU VRAM: 8.47 GB\nReserved GPU VRAM: 10.12 GB\n\nRunning trial 2 of 200...\n* Parameters:\n  * direction_index = per layer\n  * attn.o_proj.max_weight = 1.49\n  * attn.o_proj.max_weight_position = 25.36\n  * attn.o_proj.min_weight = 0.01\n  * attn.o_proj.min_weight_distance = 3.19\n  * mlp.down_proj.max_weight = 1.49\n  * mlp.down_proj.max_weight_position = 23.56\n  * mlp.down_proj.min_weight = 0.33\n  * mlp.down_proj.min_weight_distance = 11.33\n* Resetting model...\n* Abliterating...\n* Evaluating...\n  * Obtaining first-token probability distributions...\n  * KL divergence: 0.0125\n  * Counting model refusals...\n  * Refusals: 88/100\n\nElapsed time: 1m 46s\nEstimated remaining time: 2h 55m\nResident system RAM: 2.16 GB\nAllocated GPU VRAM: 8.47 GB\nReserved GPU VRAM: 10.13 GB\n\n[... 198 more trials ...]\n```\n\n### Trial selection [​](#trial-selection)\n\nAfter the optimization run is complete, you will be shown the [Pareto front](https://en.wikipedia.org/wiki/Pareto_front) of all trials, and asked to choose which combination of refusal count and [KL divergence](https://en.wikipedia.org/wiki/Kullback%E2%80%93Leibler_divergence) (a rough measure of the difference between the modified model and the original one) you want.\n\nIMPORTANT\n\nThe trial with the lowest number of refusals is not always the best choice. Any trial with a refusal count below 10 can be assumed to strongly suppress refusals, and among those, the trial with the lowest KL divergence should be chosen because it has the best chance of preserving the original model's intelligence.\n\n```\nOptimization finished!\n\nThe following trials resulted in Pareto optimal combinations of refusals\nand KL divergence. After selecting a trial, you will be able to save the\nmodel, upload it to Hugging Face, chat with it to test how well it works,\nor run standard benchmarks on it. You can return to this menu later to\nselect a different trial. Note that KL divergence values above 0.5 usually\nindicate significant damage to the original model's capabilities.\n\n? Which trial do you want to use? (Use arrow keys)\n » [Trial  91] Refusals: 11/100, KL divergence: 0.0508\n   [Trial  81] Refusals: 21/100, KL divergence: 0.0366\n   [Trial 116] Refusals: 30/100, KL divergence: 0.0296\n   [Trial  93] Refusals: 45/100, KL divergence: 0.0120\n   [Trial  70] Refusals: 74/100, KL divergence: 0.0112\n   [Trial  37] Refusals: 75/100, KL divergence: 0.0109\n   [Trial  83] Refusals: 76/100, KL divergence: 0.0077\n   [Trial  85] Refusals: 77/100, KL divergence: 0.0059\n   [Trial 126] Refusals: 80/100, KL divergence: 0.0035\n   [Trial  47] Refusals: 86/100, KL divergence: 0.0032\n   [Trial 119] Refusals: 87/100, KL divergence: 0.0011\n   [Trial 113] Refusals: 91/100, KL divergence: 0.0010\n   [Trial  49] Refusals: 93/100, KL divergence: 0.0006\n   Run additional trials\n   Exit program\n```\n\n### Action selection [​](#action-selection)\n\nAfter choosing a trial, you can now decide what you want to do with the resulting model. If you opt to upload the model to Hugging Face and you aren't already logged into Hugging Face on your system, you will be prompted to enter an access token, which you can get from the Hugging Face website as explained [here](https://huggingface.co/docs/hub/en/security-tokens). Make sure the token has write permissions, as otherwise you won't be able to upload a model with it.\n\n```\nRestoring model from trial 91...\n* Parameters:\n  * direction_index = 19.93\n  * attn.o_proj.max_weight = 1.36\n  * attn.o_proj.max_weight_position = 18.63\n  * attn.o_proj.min_weight = 1.21\n  * attn.o_proj.min_weight_distance = 16.30\n  * mlp.down_proj.max_weight = 1.44\n  * mlp.down_proj.max_weight_position = 18.86\n  * mlp.down_proj.min_weight = 1.16\n  * mlp.down_proj.min_weight_distance = 12.35\n* Resetting model...\n* Abliterating...\n\n? What do you want to do with the decensored model? (Use arrow keys)\n » Save the model to a local folder\n   Upload the model to Hugging Face\n   Chat with the model\n   Benchmark the model\n   Return to the trial selection menu\n```\n\n## Next steps [​](#next-steps)\n\nOnce you have successfully processed your first model with Heretic, it's time to explore [the many configuration parameters](/configuration) Heretic offers that give you control over what it does.\n\nIf you encounter any problems while using Heretic, or if you have suggestions or ideas for improving Heretic, please file an issue [on GitHub](https://github.com/p-e-w/heretic), or join us [on Discord](https://discord.gg/gdXc48gSyT) or [on Matrix](https://matrix.to/#/#heretic:matrix.org)!","body_html":"<h1 id=\"tutorial\">Tutorial <a href=\"#tutorial\">​</a></h1>\n<p>Welcome, friend! With Heretic, you can remove restrictions from language models, or modify them in other interesting ways.</p>\n<p>That&#39;s right, you!</p>\n<p>You don&#39;t need to be a software engineer, you don&#39;t need a machine learning PhD, you don&#39;t need to understand the intricacies of how language models work internally, and you don&#39;t need expensive hardware.</p>\n<h2 id=\"prerequisites\">Prerequisites <a href=\"#prerequisites\">​</a></h2>\n<p>Here&#39;s what you do need:</p>\n<ul><li></li></ul>\n<p>A GPU. Both Nvidia and AMD GPUs are well supported by Heretic. While Heretic also supports processing language models in CPU-only mode, this is orders of magnitude slower than accelerated processing, and should only be attempted with tiny models. You will also need to install the CUDA Toolkit if you have an Nvidia GPU, and ROCm if you have an AMD GPU.</p>\n<ul><li></li></ul>\n<p>Sufficient VRAM. Make sure that your GPU is a match for the model you want to process. As a rule of thumb, you need about 2.5 GB of VRAM per billion model parameters. So to process a model like <a href=\"https://huggingface.co/Qwen/Qwen3.5-4B\" rel=\"nofollow ugc noopener\">Qwen3.5-4B</a> (which has 4 billion parameters), you need about 10 GB of VRAM.</p>\n<ul><li></li></ul>\n<p>Python. If you run Linux, Python is almost always preinstalled, but on other platforms, you may have to <a href=\"https://www.python.org/\" rel=\"nofollow ugc noopener\">download</a> and install it manually. Heretic requires Python 3.10 or later.</p>\n<h2 id=\"set-up-a-python-virtual-environment\">Set up a Python virtual environment <a href=\"#set-up-a-python-virtual-environment\">​</a></h2>\n<p>We strongly recommend installing Heretic (and any other Python application) into its own virtual environment, which you can set up withsh</p>\n<pre><code>python -m venv venv\nsource venv/bin/activate</code></pre>\n<h2 id=\"install-pytorch\">Install PyTorch <a href=\"#install-pytorch\">​</a></h2>\n<p>TIP</p>\n<p>If you are using a cloud computing platform, the appropriate version of PyTorch is often already preinstalled.</p>\n<p>Heretic uses PyTorch to accelerate mathematical operations. Unlike all other dependencies of Heretic, PyTorch must be installed manually because the correct installation command depends on your GPU and accelerator library.</p>\n<p>You can find full installation instructions <a href=\"https://pytorch.org/get-started/locally/\" rel=\"nofollow ugc noopener\">on the PyTorch website</a>. In most cases, if you have an Nvidia GPU, you will runsh</p>\n<pre><code>pip install torch torchvision</code></pre>\n<p>and if you have an AMD GPU, you will runsh</p>\n<pre><code>pip install torch torchvision --index-url https://download.pytorch.org/whl/rocm7.2</code></pre>\n<p>but it&#39;s a good idea to check the page linked above for the specific instructions for your setup.</p>\n<h2 id=\"install-heretic\">Install Heretic <a href=\"#install-heretic\">​</a></h2>\n<p>There are <a href=\"/installation\">many ways</a> to download and install Heretic. Most users will want to runsh</p>\n<pre><code>pip install -U heretic-llm</code></pre>\n<p>and be done with it.</p>\n<h2 id=\"run-heretic\">Run Heretic <a href=\"#run-heretic\">​</a></h2>\n<p>Pick a language model <a href=\"https://huggingface.co/models?library=transformers\" rel=\"nofollow ugc noopener\">on Hugging Face</a> and copy its model ID (the part of its URL after <code>https://huggingface.co/</code>). Then runsh</p>\n<pre><code>heretic Qwen/Qwen3.5-4B</code></pre>\n<p>Replace <code>Qwen/Qwen3.5-4B</code> with the model ID you copied above.</p>\n<p>That&#39;s it! The process is fully automatic and will prompt you when you need to make a decision. Heretic does not require configuration, although <a href=\"/configuration\">many configuration parameters</a> are available, allowing you to control almost every aspect of Heretic&#39;s operation.</p>\n<p>Let&#39;s walk through what happens when Heretic processes a model:</p>\n<h3 id=\"startup\">Startup <a href=\"#startup\">​</a></h3>\n<p>At the start of the program, Heretic identifies the GPU(s) available. As you can see, even with a modest GPU (an RTX 3060 in this case) it&#39;s already possible to process small models (4B in this case).</p>\n<p>TIP</p>\n<p>Heretic supports loading models with 4-bit quantization using bitsandbytes, which can reduce the amount of VRAM required by about 70%. See the <a href=\"/configuration#quantization\"><code>quantization</code></a> setting for details.</p>\n<pre><code>█░█░█▀▀░█▀▄░█▀▀░▀█▀░█░█▀▀  v1.3.0\n█▀█░█▀▀░█▀▄░█▀▀░░█░░█░█░░\n▀░▀░▀▀▀░▀░▀░▀▀▀░░▀░░▀░▀▀▀  https://github.com/p-e-w/heretic\n\nDetected 1 CUDA device(s) (11.63 GB total VRAM)\nCUDA Version: 12.8\nDriver Version: 580.159.03\n* CUDA 0: NVIDIA GeForce RTX 3060 (11.63 GB)</code></pre>\n<h3 id=\"model-loading\">Model loading <a href=\"#model-loading\">​</a></h3>\n<p>Heretic now loads the requested model (downloading it from Hugging Face unless it is already cached), analyzes its architecture, and displays the amount of memory it occupies.</p>\n<pre><code>Loading model Qwen/Qwen3.5-4B...\n* Trying dtype auto...\n* LoRA adapters initialized (target types: down_proj, o_proj, out_proj)\n* Transformer model with 32 layers\n* Abliterable components:\n  * attn.o_proj: 32 modules total\n  * mlp.down_proj: 32 modules total\n\nResident system RAM: 1.83 GB\nAllocated GPU VRAM: 8.47 GB\nReserved GPU VRAM: 8.52 GB</code></pre>\n<h3 id=\"prompt-loading\">Prompt loading <a href=\"#prompt-loading\">​</a></h3>\n<p>The prompt datasets used for calculating refusal directions are now being loaded.</p>\n<pre><code>Loading good prompts from mlabonne/harmless_alpaca...\n* 400 prompts loaded\n\nLoading bad prompts from mlabonne/harmful_behaviors...\n* 400 prompts loaded</code></pre>\n<h3 id=\"automatic-batch-size-determination\">Automatic batch size determination <a href=\"#automatic-batch-size-determination\">​</a></h3>\n<p>Heretic does inference on prompts in batches, which dramatically speeds up processing. This step determines the largest batch size your system can support, in order to maximize processing speed.</p>\n<pre><code>Determining optimal batch size...\n* Trying batch size 1... Ok (25 tokens/s)\n* Trying batch size 2... Ok (44 tokens/s)\n* Trying batch size 4... Ok (78 tokens/s)\n* Trying batch size 8... Ok (122 tokens/s)\n* Trying batch size 16... Ok (178 tokens/s)\n* Trying batch size 32... Ok (225 tokens/s)\n* Trying batch size 64... Failed (CUDA out of memory)\n* Chosen batch size: 32</code></pre>\n<h3 id=\"automatic-response-prefix-determination\">Automatic response prefix determination <a href=\"#automatic-response-prefix-determination\">​</a></h3>\n<p>Reasoning models often start their responses with a thinking block (such as <code>&lt;think&gt;...&lt;/think&gt;</code>). Skipping the tokens in this block improves accuracy when calculating refusal directions, and Heretic automatically tests for such prefixes in order to do this.</p>\n<p>In this case, no prefix is found.</p>\n<pre><code>Checking for common response prefix...\n* None found</code></pre>\n<h3 id=\"evaluation-prompt-loading\">Evaluation prompt loading <a href=\"#evaluation-prompt-loading\">​</a></h3>\n<p>The prompt datasets used for evaluating model performance are now being loaded, and an initial evaluation is performed.</p>\n<pre><code>Loading good evaluation prompts from mlabonne/harmless_alpaca...\n* 100 prompts loaded\n* Obtaining first-token probability distributions...\n\nLoading bad evaluation prompts from mlabonne/harmful_behaviors...\n* 100 prompts loaded\n* Counting model refusals...\n* Initial refusals: 93/100</code></pre>\n<h3 id=\"refusal-directions-calculation\">Refusal directions calculation <a href=\"#refusal-directions-calculation\">​</a></h3>\n<p>Heretic modifies language models by ablating &quot;refusal directions&quot; from certain operators. This step computes those directions based on the provided prompts.</p>\n<pre><code>Calculating per-layer refusal directions...\n* Obtaining residual mean for good prompts...\n* Obtaining residual mean for bad prompts...</code></pre>\n<h3 id=\"parameter-optimization\">Parameter optimization <a href=\"#parameter-optimization\">​</a></h3>\n<p>This is the heart of the process. Heretic searches for a combination of parameter values that, when used to control ablation of the refusal directions, yields the best compromise between refusal suppression and model quality.</p>\n<p>This step can take some time, slightly under 3 hours in this case. An estimate for the remaining amount of time, as well as memory stats, are displayed after every trial.</p>\n<pre><code>Running trial 1 of 200...\n* Parameters:\n  * direction_index = per layer\n  * attn.o_proj.max_weight = 1.39\n  * attn.o_proj.max_weight_position = 25.29\n  * attn.o_proj.min_weight = 0.15\n  * attn.o_proj.min_weight_distance = 2.39\n  * mlp.down_proj.max_weight = 1.23\n  * mlp.down_proj.max_weight_position = 20.65\n  * mlp.down_proj.min_weight = 0.19\n  * mlp.down_proj.min_weight_distance = 11.58\n* Resetting model...\n* Abliterating...\n* Evaluating...\n  * Obtaining first-token probability distributions...\n  * KL divergence: 0.0126\n  * Counting model refusals...\n  * Refusals: 88/100\n\nElapsed time: 53s\nEstimated remaining time: 2h 56m\nResident system RAM: 2.23 GB\nAllocated GPU VRAM: 8.47 GB\nReserved GPU VRAM: 10.12 GB\n\nRunning trial 2 of 200...\n* Parameters:\n  * direction_index = per layer\n  * attn.o_proj.max_weight = 1.49\n  * attn.o_proj.max_weight_position = 25.36\n  * attn.o_proj.min_weight = 0.01\n  * attn.o_proj.min_weight_distance = 3.19\n  * mlp.down_proj.max_weight = 1.49\n  * mlp.down_proj.max_weight_position = 23.56\n  * mlp.down_proj.min_weight = 0.33\n  * mlp.down_proj.min_weight_distance = 11.33\n* Resetting model...\n* Abliterating...\n* Evaluating...\n  * Obtaining first-token probability distributions...\n  * KL divergence: 0.0125\n  * Counting model refusals...\n  * Refusals: 88/100\n\nElapsed time: 1m 46s\nEstimated remaining time: 2h 55m\nResident system RAM: 2.16 GB\nAllocated GPU VRAM: 8.47 GB\nReserved GPU VRAM: 10.13 GB\n\n[... 198 more trials ...]</code></pre>\n<h3 id=\"trial-selection\">Trial selection <a href=\"#trial-selection\">​</a></h3>\n<p>After the optimization run is complete, you will be shown the <a href=\"https://en.wikipedia.org/wiki/Pareto_front\" rel=\"nofollow ugc noopener\">Pareto front</a> of all trials, and asked to choose which combination of refusal count and <a href=\"https://en.wikipedia.org/wiki/Kullback%E2%80%93Leibler_divergence\" rel=\"nofollow ugc noopener\">KL divergence</a> (a rough measure of the difference between the modified model and the original one) you want.</p>\n<p>IMPORTANT</p>\n<p>The trial with the lowest number of refusals is not always the best choice. Any trial with a refusal count below 10 can be assumed to strongly suppress refusals, and among those, the trial with the lowest KL divergence should be chosen because it has the best chance of preserving the original model&#39;s intelligence.</p>\n<pre><code>Optimization finished!\n\nThe following trials resulted in Pareto optimal combinations of refusals\nand KL divergence. After selecting a trial, you will be able to save the\nmodel, upload it to Hugging Face, chat with it to test how well it works,\nor run standard benchmarks on it. You can return to this menu later to\nselect a different trial. Note that KL divergence values above 0.5 usually\nindicate significant damage to the original model&#39;s capabilities.\n\n? Which trial do you want to use? (Use arrow keys)\n » [Trial  91] Refusals: 11/100, KL divergence: 0.0508\n   [Trial  81] Refusals: 21/100, KL divergence: 0.0366\n   [Trial 116] Refusals: 30/100, KL divergence: 0.0296\n   [Trial  93] Refusals: 45/100, KL divergence: 0.0120\n   [Trial  70] Refusals: 74/100, KL divergence: 0.0112\n   [Trial  37] Refusals: 75/100, KL divergence: 0.0109\n   [Trial  83] Refusals: 76/100, KL divergence: 0.0077\n   [Trial  85] Refusals: 77/100, KL divergence: 0.0059\n   [Trial 126] Refusals: 80/100, KL divergence: 0.0035\n   [Trial  47] Refusals: 86/100, KL divergence: 0.0032\n   [Trial 119] Refusals: 87/100, KL divergence: 0.0011\n   [Trial 113] Refusals: 91/100, KL divergence: 0.0010\n   [Trial  49] Refusals: 93/100, KL divergence: 0.0006\n   Run additional trials\n   Exit program</code></pre>\n<h3 id=\"action-selection\">Action selection <a href=\"#action-selection\">​</a></h3>\n<p>After choosing a trial, you can now decide what you want to do with the resulting model. If you opt to upload the model to Hugging Face and you aren&#39;t already logged into Hugging Face on your system, you will be prompted to enter an access token, which you can get from the Hugging Face website as explained <a href=\"https://huggingface.co/docs/hub/en/security-tokens\" rel=\"nofollow ugc noopener\">here</a>. Make sure the token has write permissions, as otherwise you won&#39;t be able to upload a model with it.</p>\n<pre><code>Restoring model from trial 91...\n* Parameters:\n  * direction_index = 19.93\n  * attn.o_proj.max_weight = 1.36\n  * attn.o_proj.max_weight_position = 18.63\n  * attn.o_proj.min_weight = 1.21\n  * attn.o_proj.min_weight_distance = 16.30\n  * mlp.down_proj.max_weight = 1.44\n  * mlp.down_proj.max_weight_position = 18.86\n  * mlp.down_proj.min_weight = 1.16\n  * mlp.down_proj.min_weight_distance = 12.35\n* Resetting model...\n* Abliterating...\n\n? What do you want to do with the decensored model? (Use arrow keys)\n » Save the model to a local folder\n   Upload the model to Hugging Face\n   Chat with the model\n   Benchmark the model\n   Return to the trial selection menu</code></pre>\n<h2 id=\"next-steps\">Next steps <a href=\"#next-steps\">​</a></h2>\n<p>Once you have successfully processed your first model with Heretic, it&#39;s time to explore <a href=\"/configuration\">the many configuration parameters</a> Heretic offers that give you control over what it does.</p>\n<p>If you encounter any problems while using Heretic, or if you have suggestions or ideas for improving Heretic, please file an issue <a href=\"https://github.com/p-e-w/heretic\" rel=\"nofollow ugc noopener\">on GitHub</a>, or join us <a href=\"https://discord.gg/gdXc48gSyT\" rel=\"nofollow ugc noopener\">on Discord</a> or <a href=\"https://matrix.to/#/#heretic:matrix.org\" rel=\"nofollow ugc noopener\">on Matrix</a>!</p>","headings":[{"level":1,"text":"Tutorial ​","id":"tutorial"},{"level":2,"text":"Prerequisites ​","id":"prerequisites"},{"level":2,"text":"Set up a Python virtual environment ​","id":"set-up-a-python-virtual-environment"},{"level":2,"text":"Install PyTorch ​","id":"install-pytorch"},{"level":2,"text":"Install Heretic ​","id":"install-heretic"},{"level":2,"text":"Run Heretic ​","id":"run-heretic"},{"level":3,"text":"Startup ​","id":"startup"},{"level":3,"text":"Model loading ​","id":"model-loading"},{"level":3,"text":"Prompt loading ​","id":"prompt-loading"},{"level":3,"text":"Automatic batch size determination ​","id":"automatic-batch-size-determination"},{"level":3,"text":"Automatic response prefix determination ​","id":"automatic-response-prefix-determination"},{"level":3,"text":"Evaluation prompt loading ​","id":"evaluation-prompt-loading"},{"level":3,"text":"Refusal directions calculation ​","id":"refusal-directions-calculation"},{"level":3,"text":"Parameter optimization ​","id":"parameter-optimization"},{"level":3,"text":"Trial selection ​","id":"trial-selection"},{"level":3,"text":"Action selection ​","id":"action-selection"},{"level":2,"text":"Next steps ​","id":"next-steps"}]}}