本文摘要本文基于LM Studio,总结当前 各主流LLM本地部署的启动参数配置大全,以发挥个人算力机器的价值,文中所有的配置都经过验证可行。LuffyTheFox/qwen3.6-35b-a3b-uncensored-genesis-hermes-v6链接:https://huggingface.co/LuffyTheFox/Qwen3.6-35B-A3B-Uncensored-Genesis-Herm...

本文基于LM Studio,总结当前 各主流LLM本地部署的启动参数配置大全,以发挥个人算力机器的价值,文中所有的配置都经过验证可行。
LuffyTheFox/qwen3.6-35b-a3b-uncensored-genesis-hermes-v6
链接:https://huggingface.co/LuffyTheFox/Qwen3.6-35B-A3B-Uncensored-Genesis-Hermes-V6-GGUF
启动参数配置:
| 类型 | 参数项 | 参数值 |
|---|---|---|
| Model Information | ||
| Model name | LuffyTheFox/qwen3.6-35b-a3b-uncensored-genesis-hermes-v6 | |
| Format | GGUF | |
| Quantization | Q8_0 | |
| Arch | qwen35moe | |
| Capabilities | V,T,R | |
| Domain | llm | |
| Size on disk | 37.80G | |
| Context and Offload | ||
| Context Length | 65536 | |
| GPU Offload | 40 | |
| Advanced | ||
| CPU Thread Pool Size | 16 | |
| Evaluation Batch Size | 8192 | |
| Physical Batch Size | 512 | |
| Max Concurr.. | 2 | |
| Unified KV Cache | 开 | |
| Context Checkpoints | 32 | |
| Reasoning Budget Message | 空 | |
| RoPE Frequency Base | 不选,Auto | |
| RoPE Frequency Scale | 不选,Auto | |
| Offload KV Cache to GPU Memory | 开 | |
| keep Model in Memory | 开 | |
| Try mmap() | 关 | |
| Seed | 不选,Random Seed | |
| Number of Experts | 8 | |
| Number of la... | 16 | |
| Speculative Decoding | Off | |
| Chat Template | 默认 | |
| Flash Attention | 关 | |
| K Cache Quantization Ty... | 不选,Experimental | |
| V Cache Quantization Ty... | 不选,Experimental | |
| Inference | System Prompt | 默认 |
| Reasoning | Enable Thinking | 开 |
| Reasoning Budget | 不选,Unrestricted | |
| Setting | Temperature | 0.1 |
| Limit Response Length | 不选 | |
| Context Overflow | Truncate Middle | |
| Stop Stri... | 默认 | |
| Sampling | Top K Sampling | 40 |
| Repeat Penalty | 选中,1.1 | |
| Presence Penalty | 不选 | |
| Top P Sampling | 选中,0.95 | |
| Structuerd Output | Structuerd Output | 关 |
应用效果:

觉得内容不错?我要