Skip to main content
After Heretic successfully decensors a model, you can upload it to Hugging Face Hub to share with the community or deploy it to production. This guide covers the complete upload workflow.

Post-Processing Workflow

Once optimization is complete, Heretic presents you with options:
1

Select a Trial

Choose from the Pareto-optimal trials based on your refusal/quality tradeoff preference
2

Choose Action

Select “Upload the model to Hugging Face” from the action menu
3

Authenticate

Provide your Hugging Face access token when prompted
4

Configure Upload

Set repository name and visibility
5

Upload Complete

Model is pushed to Hugging Face Hub with auto-generated model card

Authentication

Heretic needs a Hugging Face access token to upload models.

Using Existing Token

If you’ve already logged in via huggingface-cli:
Heretic will automatically detect and use your stored token.

Providing Token Manually

If no token is found, Heretic will prompt you:
To create a token:
  1. Visit https://huggingface.co/settings/tokens
  2. Click “New token”
  3. Select “Write” permissions
  4. Copy the token and paste when prompted
Heretic does NOT store the token to disk for security reasons. You’ll need to re-enter it if you restart the program.

Token Verification

After providing a token, Heretic confirms your identity:

Repository Configuration

Repository Name

Heretic suggests a default name following best practices:
Default format: {username}/{original-model-name}-heretic Examples:
  • Original: Qwen/Qwen3-4B-Instruct-2507
  • Suggested: username/Qwen3-4B-Instruct-2507-heretic
The -heretic suffix helps users identify decensored models and is recognized by the community.

Visibility

Choose whether your model should be public or private:

Public

Visible to everyone, appears in search results, contributes to the community

Private

Only visible to you and collaborators, useful for testing or proprietary models

Upload Process

Merged Model vs LoRA Adapter

Heretic gives you a choice on what to upload:

Quantized Model Warning

If you loaded the model with quantization, merging requires additional RAM:
RAM Requirements for Merging:
  • Rule of thumb: ~3x the parameter count in GB
  • 27B model: ~80 GB RAM
  • 70B model: ~200 GB RAM
If you don’t have enough RAM, choose “Cancel” or save as LoRA adapter only.
See Quantization - Merging Quantized Models for details.

Model Card Generation

Heretic automatically generates a comprehensive model card:

Auto-Generated Content

The model card includes:
1

Introduction Section

Description of the decensoring process and Heretic version used
2

Performance Metrics

Refusal rates and KL divergence for the selected trial
3

Trial Parameters

Complete parameter configuration for reproducibility
4

Tags

Automatic tags for discoverability:
  • heretic
  • uncensored
  • decensored
  • abliterated

Preserved Original Content

If the original model has a README:
  • Original content is preserved
  • Heretic introduction is prepended
  • Original tags are kept (+ new tags added)
  • Model architecture info retained
The generated model card helps users understand how your model was created and sets expectations for its behavior.

Naming Conventions

The Heretic community has established naming conventions:

Standard Format

Examples:
  • p-e-w/gemma-3-12b-it-heretic
  • p-e-w/gpt-oss-20b-heretic
  • p-e-w/Qwen3-4B-Instruct-2507-heretic

Why Use the Suffix?

Recognition

Users can instantly identify Heretic-processed models

Community

Join 1000+ other Heretic models on the Hub

Consistency

Follows established community standards

Community Models

The Heretic community has created and published over 1,000 models:

Browse All Heretic Models

Visit the Hugging Face Hub:

The Bestiary Collection

Curated collection of high-quality Heretic models created by the project maintainer:
Includes models like:
  • p-e-w/gemma-3-12b-it-heretic
  • p-e-w/gpt-oss-20b-heretic
  • p-e-w/Qwen3-4B-Instruct-2507-heretic
Browse The Bestiary for examples of well-configured Heretic models and inspiration for your own uploads.

Upload Workflow Example

Complete example of uploading a model:
Your model is now available at:

Best Practices

1

Test Before Uploading

Use the “Chat with the model” option to verify quality
2

Choose the Right Trial

Balance refusal suppression vs KL divergence for your use case
  • Low KL divergence (less than 0.5): Better preserves original capabilities
  • Low refusals (less than 5/100): More effective decensoring
3

Use Descriptive Names

Include the base model name and -heretic suffix✅ username/llama-3.1-8b-instruct-heretic
❌ username/my-uncensored-model
4

Set Appropriate Visibility

Start with private for testing, make public when satisfied
5

Add Custom README Content

Edit the model card after upload to add:
  • Usage examples
  • Benchmark results
  • Known limitations
  • License information

Troubleshooting

Authentication Failed

Error: Invalid token or permission denied Solutions:

Upload Failed

Error: Network error or timeout during upload Solutions:
  • Check internet connection
  • Try uploading during off-peak hours
  • Save locally first, then upload manually:

Insufficient RAM for Merge

Error: System freezes or OOM during merge Solutions:
  1. Save LoRA adapter only:
  2. Merge on a larger machine:
  3. Use cloud instance:
    • Rent a high-RAM instance temporarily
    • Load model, merge, and upload from there

Local Save Option

Before or instead of uploading, you can save locally:
This is useful for:
  • Testing before upload
  • Offline deployment
  • Manual upload later via huggingface-cli