AI SeedbankHelp preserve open and free AI for humanity's future

← All models

Qwen_Qwen3-4B-Instruct-2507

Qwen · View on Hugging Face ↗

Get this model

Download TorrentMagnet Link

Seeders: 1 · Leechers: 0

Observed 2026-09-02T13:56:39Z via announce.aitorrent.org:7070.

Model card

The complete upstream card, rendered from this payload's README.md — the same hash-verified bytes the torrent distributes. Images and off-site links are removed; the original card on Hugging Face carries them.


library_name: transformers license: apache-2.0 license_link: https://huggingface.co/Qwen/Qwen3-4B-Instruct-2507/blob/main/LICENSE pipeline_tag: text-generation

Qwen3-4B-Instruct-2507

Highlights

We introduce the updated version of the Qwen3-4B non-thinking mode, named Qwen3-4B-Instruct-2507, featuring the following key enhancements:

  • Significant improvements in general capabilities, including instruction following, logical reasoning, text comprehension, mathematics, science, coding and tool usage.
  • Substantial gains in long-tail knowledge coverage across multiple languages.
  • Markedly better alignment with user preferences in subjective and open-ended tasks, enabling more helpful responses and higher-quality text generation.
  • Enhanced capabilities in 256K long-context understanding.

Model Overview

Qwen3-4B-Instruct-2507 has the following features:

  • Type: Causal Language Models
  • Training Stage: Pretraining & Post-training
  • Number of Parameters: 4.0B
  • Number of Paramaters (Non-Embedding): 3.6B
  • Number of Layers: 36
  • Number of Attention Heads (GQA): 32 for Q and 8 for KV
  • Context Length: 262,144 natively.

NOTE: This model supports only non-thinking mode and does not generate <think></think> blocks in its output. Meanwhile, specifying enable_thinking=False is no longer required.

For more details, including benchmark evaluation, hardware requirements, and inference performance, please refer to our blog, GitHub, and Documentation.

Performance

GPT-4.1-nano-2025-04-14 Qwen3-30B-A3B Non-Thinking Qwen3-4B Non-Thinking Qwen3-4B-Instruct-2507
Knowledge
MMLU-Pro 62.8 69.1 58.0 69.6
MMLU-Redux 80.2 84.1 77.3 84.2
GPQA 50.3 54.8 41.7 62.0
SuperGPQA 32.2 42.2 32.0 42.8
Reasoning
AIME25 22.7 21.6 19.1 47.4
HMMT25 9.7 12.0 12.1 31.0
ZebraLogic 14.8 33.2 35.2 80.2
LiveBench 20241125 41.5 59.4 48.4 63.0
Coding
LiveCodeBench v6 (25.02-25.05) 31.5 29.0 26.4 35.1
MultiPL-E 76.3 74.6 66.6 76.8
Aider-Polyglot 9.8 24.4 13.8 12.9
Alignment
IFEval 74.5 83.7 81.2 83.4
Arena-Hard v2* 15.9 24.8 9.5 43.4
Creative Writing v3 72.7 68.1 53.6 83.5
WritingBench 66.9 72.2 68.5 83.4
Agent
BFCL-v3 53.0 58.6 57.6 61.9
TAU1-Retail 23.5 38.3 24.3 48.7
TAU1-Airline 14.0 18.0 16.0 32.0
TAU2-Retail - 31.6 28.1 40.4
TAU2-Airline - 18.0 12.0 24.0
TAU2-Telecom - 18.4 17.5 13.2
Multilingualism
MultiIF 60.7 70.8 61.3 69.0
MMLU-ProX 56.2 65.1 49.6 61.6
INCLUDE 58.6 67.8 53.8 60.1
PolyMATH 15.6 23.3 16.6 31.1

*: For reproducibility, we report the win rates evaluated by GPT-4.1.

Quickstart

The code of Qwen3 has been in the latest Hugging Face transformers and we advise you to use the latest version of transformers.

With transformers<4.51.0, you will encounter the following error:

KeyError: 'qwen3'

The following contains a code snippet illustrating how to use the model generate content based on given inputs.

from transformers import AutoModelForCausalLM, AutoTokenizer

model_name = "Qwen/Qwen3-4B-Instruct-2507"

# load the tokenizer and the model
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(
    model_name,
    torch_dtype="auto",
    device_map="auto"
)

# prepare the model input
prompt = "Give me a short introduction to large language model."
messages = [
    {"role": "user", "content": prompt}
]
text = tokenizer.apply_chat_template(
    messages,
    tokenize=False,
    add_generation_prompt=True,
)
model_inputs = tokenizer([text], return_tensors="pt").to(model.device)

# conduct text completion
generated_ids = model.generate(
    **model_inputs,
    max_new_tokens=16384
)
output_ids = generated_ids[0][len(model_inputs.input_ids[0]):].tolist() 

content = tokenizer.decode(output_ids, skip_special_tokens=True)

print("content:", content)

For deployment, you can use sglang>=0.4.6.post1 or vllm>=0.8.5 or to create an OpenAI-compatible API endpoint:

  • SGLang:
    python -m sglang.launch_server --model-path Qwen/Qwen3-4B-Instruct-2507 --context-length 262144
    
  • vLLM:
    vllm serve Qwen/Qwen3-4B-Instruct-2507 --max-model-len 262144
    

Note: If you encounter out-of-memory (OOM) issues, consider reducing the context length to a shorter value, such as 32,768.

For local use, applications such as Ollama, LMStudio, MLX-LM, llama.cpp, and KTransformers have also supported Qwen3.

Agentic Use

Qwen3 excels in tool calling capabilities. We recommend using Qwen-Agent to make the best use of agentic ability of Qwen3. Qwen-Agent encapsulates tool-calling templates and tool-calling parsers internally, greatly reducing coding complexity.

To define the available tools, you can use the MCP configuration file, use the integrated tool of Qwen-Agent, or integrate other tools by yourself.

from qwen_agent.agents import Assistant

# Define LLM
llm_cfg = {
    'model': 'Qwen3-4B-Instruct-2507',

    # Use a custom endpoint compatible with OpenAI API:
    'model_server': 'http://localhost:8000/v1',  # api_base
    'api_key': 'EMPTY',
}

# Define Tools
tools = [
    {'mcpServers': {  # You can specify the MCP configuration file
            'time': {
                'command': 'uvx',
                'args': ['mcp-server-time', '--local-timezone=Asia/Shanghai']
            },
            "fetch": {
                "command": "uvx",
                "args": ["mcp-server-fetch"]
            }
        }
    },
  'code_interpreter',  # Built-in tools
]

# Define Agent
bot = Assistant(llm=llm_cfg, function_list=tools)

# Streaming generation
messages = [{'role': 'user', 'content': 'https://qwenlm.github.io/blog/ Introduce the latest developments of Qwen'}]
for responses in bot.run(messages=messages):
    pass
print(responses)

Best Practices

To achieve optimal performance, we recommend the following settings:

  1. Sampling Parameters:

    • We suggest using Temperature=0.7, TopP=0.8, TopK=20, and MinP=0.
    • For supported frameworks, you can adjust the presence_penalty parameter between 0 and 2 to reduce endless repetitions. However, using a higher value may occasionally result in language mixing and a slight decrease in model performance.
  2. Adequate Output Length: We recommend using an output length of 16,384 tokens for most queries, which is adequate for instruct models.

  3. Standardize Output Format: We recommend using prompts to standardize model outputs when benchmarking.

    • Math Problems: Include "Please reason step by step, and put your final answer within \boxed{}." in the prompt.
    • Multiple-Choice Questions: Add the following JSON structure to the prompt to standardize responses: "Please show your choice in the answer field with only the choice letter, e.g., "answer": "C"."

Citation

If you find our work helpful, feel free to give us a cite.

@misc{qwen3technicalreport,
      title={Qwen3 Technical Report}, 
      author={Qwen Team},
      year={2025},
      eprint={2505.09388},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2505.09388}, 
}

Magnet link

Opens the swarm directly in your torrent client — no file download needed. Copy-paste works too:

magnet:?xt=urn:btih:5efc073924f84dbacea29b396d6683c0c0591624&dn=Qwen_Qwen3-4B-Instruct-2507

Open magnet in torrent client · infohash 5efc073924f84dbacea29b396d6683c0c0591624

Files & hashes

PathSizesha1sha256
LICENSE11.1 KB (11,343 B)6634c8cc3133b3848ec74b9f275acaaa1ea618ab832dd9e00a68dd83b3c3fb9f5588dad7dcf337a0db50f7d9483f310cd292e92e
README.md8.0 KB (8,168 B)b85879f8e465b5b02bd7ea13dd7ffaf22816bf598e3dd0c3b5b11897cc71092ccfe517bb7a9783479baa3665aad73c8d1a2041cd
config.json727 B (727 B)6988f134db143052042f2bd6e0c897bc6a6051895beea1a4a34c62782bfb2f911c606741a3bab8f92d80a118fa053c28af12e8ba
generation_config.json238 B (238 B)432531a002c181a19de338313d2375e9d7494d7e835fffe355c9438e7a25be099b3fccaa98350b83451f9fd2d99512e74f1ade48
merges.txt1.6 MB (1,671,839 B)20024bfe7c83998e9aeaf98a0cd6a2ce6306c2f0599bab54075088774b1733fde865d5bd747cbcc7a547c5bc12610e874e26f5e3
model-00001-of-00003.safetensors3.69 GB (3,957,900,840 B)74aebba0804ab8c79df005708fd94faac88f106e75311d91bb08cf0b882913da464a1e722a31fb44db35208663487efb7a3d8ed6
model-00002-of-00003.safetensors3.71 GB (3,987,450,520 B)032b8669af31e5d6844c39617d3da4369eaeed160b48adbb1f60e901153d91907ba11ce63bd4b8b584482e730f48808d055dfba1
model-00003-of-00003.safetensors95.0 MB (99,630,640 B)32c9254c7ae0255a80327de2b9be8041a00237a67dd39ccca5e4de123c74c14af44c9bf2eb75df33b4614382af0134528e060d5d
model.safetensors.index.json32.0 KB (32,819 B)4747b0297d3109f14db49886972e3369c9a00b2ad6c42883a895dfef5b0080ed2116a1bcd764f558406b98923d675978a1abf29c
tokenizer.json10.9 MB (11,422,654 B)a1de58e2833d77bb504a8e430b1b25d359912a98aeb13307a71acd8fe81861d94ad54ab689df773318809eed3cbe794b4492dae4
tokenizer_config.json9.2 KB (9,377 B)51c1be0d9192e7f6e6596de71d0f07d58fbc32aca62ff0a2472a0fa1b8eaabcb57c59b58afa42a22831dc141400b6e0cf2b65ce3
vocab.json2.6 MB (2,776,833 B)4783fe10ac3adce15ac8f358ef5462739852c569ca10d7e9fb3ed18575dd1e277a2579c16d108e32f27439684afa0e10b1440910

Cite this release

Canonical URL
https://aiseedbank.org/models/Qwen_Qwen3-4B-Instruct-2507/
Slug
Qwen_Qwen3-4B-Instruct-2507
Infohash
5efc073924f84dbacea29b396d6683c0c0591624
License
apache-2.0
Signing key fingerprint
85a3b32c3712427b

Every file carries a locally computed sha256 — verify a download against the signed sums: Qwen_Qwen3-4B-Instruct-2507.SHA256SUMS (+ minisign signature).

Provenance

Upstream repositoryQwen/Qwen3-4B-Instruct-2507
Revision (pinned)cdbee75f17c01a7cc42f958dc650907174af0554
Fetched at2026-09-02T05:41:13Z
License at fetchapache-2.0
Snapshot toolhuggingface · seedbank 0.1.0

Trackers

✓ verified · rehash-vs-hf-metadata at 2026-09-02T05:42:26Z

apache-2.07.51 GB (8,060,915,998 bytes)transformerssafetensorsqwen3text-generationconversationaleval-resultstext-generation-inferenceendpoints_compatiblepaper: 2505.09388