Help preserve open and free AI for humanity's future

← All models

Qwen_Qwen3-Coder-30B-A3B-Instruct-FP8

Qwen · View on Hugging Face ↗

Qwen3-Coder 30B-A3B instruct (sparse MoE, FP8 weights) for code generation and agentic coding.

✓ verified · rehash-vs-hf-metadata at 2026-08-22T22:47:12Z

apache-2.029.05 GB (31,195,131,256 bytes)transformerssafetensorsqwen3_moetext-generationconversationalendpoints_compatiblefp8paper: 2505.09388

Get this model

Download Qwen_Qwen3-Coder-30B-A3B-Instruct-FP8.torrent

Recommended — the .torrent carries the webseed url-list, so your client can fall back to plain HTTPS if the swarm is thin. See/verify for the full download + verification walkthrough.

Model card

The complete upstream card, rendered from this payload's README.md — the same hash-verified bytes the torrent distributes. Images and off-site links are removed; the original card on Hugging Face carries them.


library_name: transformers license: apache-2.0 license_link: https://huggingface.co/Qwen/Qwen3-Coder-30B-A3B-Instruct-FP8/blob/main/LICENSE pipeline_tag: text-generation

Qwen3-Coder-30B-A3B-Instruct-FP8

Highlights

Qwen3-Coder is available in multiple sizes. Today, we're excited to introduce Qwen3-Coder-30B-A3B-Instruct-FP8. This streamlined model maintains impressive performance and efficiency, featuring the following key enhancements:

  • Significant Performance among open models on Agentic Coding, Agentic Browser-Use, and other foundational coding tasks.
  • Long-context Capabilities with native support for 256K tokens, extendable up to 1M tokens using Yarn, optimized for repository-scale understanding.
  • Agentic Coding supporting for most platform such as Qwen Code, CLINE, featuring a specially designed function call format.

Model Overview

Qwen3-Coder-30B-A3B-Instruct-FP8 has the following features:

  • Type: Causal Language Models
  • Training Stage: Pretraining & Post-training
  • Number of Parameters: 30.5B in total and 3.3B activated
  • Number of Layers: 48
  • Number of Attention Heads (GQA): 32 for Q and 4 for KV
  • Number of Experts: 128
  • Number of Activated Experts: 8
  • Context Length: 262,144 natively.

NOTE: This model supports only non-thinking mode and does not generate <think></think> blocks in its output. Meanwhile, specifying enable_thinking=False is no longer required.

For more details, including benchmark evaluation, hardware requirements, and inference performance, please refer to our blog, GitHub, and Documentation.

Quickstart

We advise you to use the latest version of transformers.

With transformers<4.51.0, you will encounter the following error:

KeyError: 'qwen3_moe'

The following contains a code snippet illustrating how to use the model generate content based on given inputs.

from transformers import AutoModelForCausalLM, AutoTokenizer

model_name = "Qwen/Qwen3-Coder-30B-A3B-Instruct-FP8"

# load the tokenizer and the model
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(
    model_name,
    torch_dtype="auto",
    device_map="auto"
)

# prepare the model input
prompt = "Write a quick sort algorithm."
messages = [
    {"role": "user", "content": prompt}
]
text = tokenizer.apply_chat_template(
    messages,
    tokenize=False,
    add_generation_prompt=True,
)
model_inputs = tokenizer([text], return_tensors="pt").to(model.device)

# conduct text completion
generated_ids = model.generate(
    **model_inputs,
    max_new_tokens=65536
)
output_ids = generated_ids[0][len(model_inputs.input_ids[0]):].tolist() 

content = tokenizer.decode(output_ids, skip_special_tokens=True)

print("content:", content)

Note: If you encounter out-of-memory (OOM) issues, consider reducing the context length to a shorter value, such as 32,768.

Note on FP8

For convenience and performance, we have provided fp8-quantized model checkpoint for Qwen3, whose name ends with -FP8. The quantization method is fine-grained fp8 quantization with block size of 128. You can find more details in the quantization_config field in config.json.

You can use the Qwen3-30B-A3B-Instruct-FP8 model with serveral inference frameworks, including transformers, sglang, and vllm, as the original bfloat16 model. However, please pay attention to the following known issues:

  • transformers:
    • there are currently issues with the "fine-grained fp8" method in transformers for distributed inference. You may need to set the environment variable CUDA_LAUNCH_BLOCKING=1 if multiple devices are used in inference.

Agentic Coding

Qwen3-Coder excels in tool calling capabilities.

You can simply define or use any tools as following example.

# Your tool implementation
def square_the_number(num: float) -> dict:
    return num ** 2

# Define Tools
tools=[
    {
        "type":"function",
        "function":{
            "name": "square_the_number",
            "description": "output the square of the number.",
            "parameters": {
                "type": "object",
                "required": ["input_num"],
                "properties": {
                    'input_num': {
                        'type': 'number', 
                        'description': 'input_num is a number that will be squared'
                        }
                },
            }
        }
    }
]

import OpenAI
# Define LLM
client = OpenAI(
    # Use a custom endpoint compatible with OpenAI API
    base_url='http://localhost:8000/v1',  # api_base
    api_key="EMPTY"
)
 
messages = [{'role': 'user', 'content': 'square the number 1024'}]

completion = client.chat.completions.create(
    messages=messages,
    model="Qwen3-Coder-30B-A3B-Instruct-FP8",
    max_tokens=65536,
    tools=tools,
)

print(completion.choice[0])

Best Practices

To achieve optimal performance, we recommend the following settings:

  1. Sampling Parameters:

    • We suggest using temperature=0.7, top_p=0.8, top_k=20, repetition_penalty=1.05.
  2. Adequate Output Length: We recommend using an output length of 65,536 tokens for most queries, which is adequate for instruct models.

Citation

If you find our work helpful, feel free to give us a cite.

@misc{qwen3technicalreport,
      title={Qwen3 Technical Report}, 
      author={Qwen Team},
      year={2025},
      eprint={2505.09388},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2505.09388}, 
}

Magnet link (secondary — no webseeds)

Opens the swarm directly, but carries no webseed url-list. Prefer the.torrent download above — HTTP fallback seeds ride inside it.

magnet:?xt=urn:btih:6b9e8a0ad05bc73dc36ad7fb4ff700ba5da86524&dn=Qwen_Qwen3-Coder-30B-A3B-Instruct-FP8

Open magnet in torrent client · infohash 6b9e8a0ad05bc73dc36ad7fb4ff700ba5da86524

Files & hashes

PathSizeMethodHash
LICENSE11.1 KB (11,343 B)sha1-git-blob6634c8cc3133b3848ec74b9f275acaaa1ea618ab
README.md6.0 KB (6,104 B)sha1-git-blobd976b9fd27abf52b88733e77586decdf9b680f02
chat_template.jinja6.1 KB (6,211 B)sha1-git-blob49b0e8d0ee7e655434942233b4b0ab2104d5247e
config.json7.0 KB (7,156 B)sha1-git-blobc942d664f161ce97e8636b51deb0b903e0f3cd0d
generation_config.json180 B (180 B)sha1-git-blobba3e8d9206508daab666e1aaa9dcc9c435bfa541
merges.txt1.6 MB (1,671,853 B)sha1-git-blob31349551d90c7606f325fe0f11bbb8bd5fa0d7c7
model-00001-of-00004.safetensors9.31 GB (10,001,462,368 B)sha256-lfs04665458e2332cc10409cd51298c95627000ed44cf29134663203e3f15c3c053
model-00002-of-00004.safetensors9.31 GB (10,000,577,408 B)sha256-lfs3d2228cdb569b07b2eb18b342c98ffb4c425264a5324446014046d402c80c1a2
model-00003-of-00004.safetensors9.31 GB (10,000,577,448 B)sha256-lfs7a6b5fd146a07c319f475589d08dd8d583b1e0919905c73a1e5af745244fc1db
model-00004-of-00004.safetensors1.09 GB (1,173,001,360 B)sha256-lfs6438b5403255c6fbf02bd20f9c75985e8a98ebc46f9f45899283ea374c25383a
model.safetensors.index.json3.4 MB (3,565,670 B)sha1-git-blobda96931fa6e4dcf2054f39faa39fd439c78586aa
qwen3coder_tool_parser.py30.9 KB (31,613 B)sha1-git-blob0960292aa3d83ff629f65be6c4bcb62632e6198d
tokenizer.json10.9 MB (11,422,654 B)sha256-lfsaeb13307a71acd8fe81861d94ad54ab689df773318809eed3cbe794b4492dae4
tokenizer_config.json12.7 KB (13,055 B)sha1-git-blobcc8cf0649efea42a8c099d852b1525fcfc42c222
vocab.json2.6 MB (2,776,833 B)sha1-git-blob4783fe10ac3adce15ac8f358ef5462739852c569

Provenance

Upstream repositoryQwen/Qwen3-Coder-30B-A3B-Instruct-FP8
Revision (pinned)dcaee4d4dfc5ee71ad501f01f530e5652438fde0
Fetched at2026-08-22T22:25:58Z
License at fetchapache-2.0
Snapshot toolhuggingface · seedbank 0.1.0

Trackers

Webseeds