# vllm > vLLM is a fast and easy-to-use library for LLM inference and serving. Use this skill when working with vLLM deployment, model serving, inference optimization, PagedAttention, continuous batching, tensor parallelism, or high-throughput LLM serving. Use when: deploying LLM inference services, optimizing throughput, implementing production LLM serving, distributed inference, or when user mentions vLLM, inference, serving, PagedAttention, 推理, 部署, tensor parallel, quantization, high-throughput. Triggers: "vLLM", "inference", "serving", "PagedAttention", "推理", "部署", "tensor parallel", "continuous batching", "quantization", "high-throughput", "模型部署", "LLM serving" - Author: wangsc522 - Repository: wangsc2024/skills - Version: 20260119111409 - Stars: 0 - Forks: 0 - Last Updated: 2026-02-06 - Source: https://github.com/wangsc2024/skills - Web: https://mule.run/skillshub/@@wangsc2024/skills~vllm:20260119111409 --- --- name: vllm description: | vLLM is a fast and easy-to-use library for LLM inference and serving. Use this skill when working with vLLM deployment, model serving, inference optimization, PagedAttention, continuous batching, tensor parallelism, or high-throughput LLM serving. Use when: deploying LLM inference services, optimizing throughput, implementing production LLM serving, distributed inference, or when user mentions vLLM, inference, serving, PagedAttention, 推理, 部署, tensor parallel, quantization, high-throughput. Triggers: "vLLM", "inference", "serving", "PagedAttention", "推理", "部署", "tensor parallel", "continuous batching", "quantization", "high-throughput", "模型部署", "LLM serving" --- # Vllm Skill Comprehensive assistance with vllm development, generated from official documentation. ## When to Use This Skill This skill should be triggered when: - Working with vllm - Asking about vllm features or APIs - Implementing vllm solutions - Debugging vllm code - Learning vllm best practices ## Quick Reference ### Common Patterns **Pattern 1:** vLLM GitHub Home User Guide User Guide Getting Started Getting Started Quickstart Installation Installation GPU CPU TPU Examples Examples Offline inference Offline inference Async LLM Streaming Audio Language Automatic Prefix Caching Basic Batch LLM Inference Chat With Tools Context Extension Data Parallel Disaggregated Prefill V1 Disaggregated Prefill Encoder Decoder Multimodal KV Load Failure Recovery Test LLM Engine Example LLM Engine Reset Kv Load Sharded State Logits Processor LoRA With Quantization Inference Metrics Mistral-Small MLPSpeculator MultiLoRA Inference Offline Inference with the OpenAI Batch file format Prefix Caching Prompt Embed Inference Qwen2.5-Omni Offline Inference Examples Qwen3 Omni Qwen 1M Reproducibility RLHF RLHF Colocate RLHF Online Quant RLHF Utils Save Sharded State Simple Profiling Skip Loading Weights In Engine Init Spec Decode Structured Outputs Torchrun Dp Example Torchrun Example Vision Language Vision Language Multi Image Online serving Online serving API Client Helm Charts Monitoring Dashboards Disaggregated Encoder Disaggregated Prefill Disaggregated Serving Disaggregated Serving P2P Nccl Xpyd Elastic Ep Gradio OpenAI Chatbot Webserver Gradio Webserver Kv Events Subscriber Multi-Node-Serving Multi Instance Data Parallel OpenAI Chat Completion Client OpenAI Chat Completion Client For Multimodal OpenAI Chat Completion Client With Tools OpenAI Chat Completion Client With Tools Required OpenAI Chat Completion Client With Tools Xlam OpenAI Chat Completion Client With Tools Xlam Streaming OpenAI Chat Completion Tool Calls With Reasoning OpenAI Chat Completion With Reasoning OpenAI Chat Completion With Reasoning Streaming OpenAI Completion Client OpenAI Responses Client OpenAI Responses Client With Mcp Tools OpenAI Responses Client With Tools OpenAI Transcription Client OpenAI Translation Client Setup OpenTelemetry POC Prometheus and Grafana Prompt Embed Inference With OpenAI Client Ray Serve Deepseek Retrieval Augmented Generation With Langchain Retrieval Augmented Generation With Llamaindex Run Cluster Sagemaker-Entrypoint Streamlit OpenAI Chatbot Webserver Structured Outputs Token Generation Client Utils Others Others LMCache Examples Logging Configuration Tensorize vLLM Model Pooling Pooling Classify Embed Plugin Pooling Score Token Classify Token Embed General General vLLM V1 Frequently Asked Questions Production Metrics Reproducibility Security Troubleshooting Usage Stats Collection Inference and Serving Inference and Serving Offline Inference OpenAI-Compatible Server Context Parallel Deployment Data Parallel Deployment Troubleshooting distributed deployments Expert Parallel Deployment Parallelism and Scaling Integrations Integrations LangChain LlamaIndex Deployment Deployment Using Docker Using Kubernetes Using Nginx Frameworks Frameworks Anyscale AnythingLLM AutoGen BentoML Cerebrium Chatbox Dify dstack Haystack Helm Hugging Face Inference Endpoints LiteLLM Lobe Chat LWS Modal Open WebUI Retrieval-Augmented Generation SkyPilot Streamlit NVIDIA Triton Integrations Integrations KAITO KServe Kthena KubeAI KubeRay Llama Stack llm-d llmaz Production stack Training Training Reinforcement Learning from Human Feedback Transformers Reinforcement Learning Configuration Configuration Conserving Memory Engine Arguments Environment Variables Model Resolution Optimization and Tuning Server Arguments TPU Models Models Supported Models Generative Models Pooling Models Extensions Extensions Loading Model weights with fastsafetensors Loading models with Run:ai Model Streamer Loading models with CoreWeave's Tensorizer Hardware Supported Models Hardware Supported Models CPU - Intel® Xeon® XPU - Intel® GPUs TPU Features Features Automatic Prefix Caching Batch Invariance Custom Arguments Custom Logits Processors Disaggregated Encoder Disaggregated Prefilling (experimental) Interleaved Thinking LoRA Adapters MooncakeConnector Usage Guide Multimodal Inputs NixlConnector Usage Guide Prompt Embedding Inputs Reasoning Outputs Sleep Mode Speculative Decoding Structured Outputs Tool Calling Quantization Quantization AutoAWQ AutoRound BitBLAS BitsAndBytes FP8 W8A8 GGUF GPTQModel FP8 INC INT4 W4A16 INT8 W8A8 NVIDIA Model Optimizer Quantized KV Cache AMD Quark AMD Quark Table of contents Quark Installation Quantization Process 1. Load the Model 2. Prepare the Calibration Dataloader 3. Set the Quantization Configuration 4. Quantize the Model and Export 5. Evaluation in vLLM Quark Quantization Script Using OCP MX (MXFP4, MXFP6) models Using Quark Quantized layerwise Auto Mixed Precision (AMP) Models 1. Quantize a model using mixed precision in AMD Quark 2. inference the quantized mixed precision model in vLLM TorchAO Developer Guide Developer Guide General General Deprecation Policy Dockerfile Incremental Compilation Workflow Profiling vLLM Vulnerability Management Model Implementation Model Implementation Basic Model Registering a Model Unit Testing Multi-Modal Support Speech-to-Text (Transcription/Translation) Support CI CI CI Failures Nightly Builds of vLLM Wheels Update PyTorch version on vLLM OSS CI/CD Design Documents Design Documents Plugins Plugins IO Processor Plugins LoRA Resolver Plugins Plugin System Architecture Overview CUDA Graphs Dual Batch Overlap How to debug the vLLM-torch.compile integration Fused MoE Modular Kernel Integration with Hugging Face Hybrid KV Cache Manager Logits Processors Metrics Multi-Modal Data Processing Fused MoE Kernel Features Python Multiprocessing Optimization levels P2P NCCL Connector Paged Attention Automatic Prefix Caching torch.compile integration Benchmarking Benchmarking Benchmark CLI Parameter Sweeps Performance Dashboard API Reference API Reference vllm vllm beam_search collect_env connections env_override envs forward_context logger logits_process logprobs outputs pooling_params sampling_params scalar_type scripts sequence tasks tracing version assets assets audio base image video attention attention layer selector backends backends abstract registry utils layers layers chunked_local_attention cross_attention encoder_only_attention mm_encoder_attention ops ops chunked_prefill_paged_decode common flashmla merge_attn_states paged_attn pallas_kv_cache_update prefix_prefill rocm_aiter_mla_sparse triton_decode_attention triton_merge_attn_states triton_reshape_and_cache_flash triton_unified_attention vit_attn_wrappers utils utils fa_utils kv_sharing_utils kv_transfer_utils benchmarks benchmarks datasets latency serve startup throughput lib lib endpoint_request_func ready_checker utils sweep sweep cli param_sweep plot plot_pareto serve serve_sla server sla_sweep utils compilation compilation activation_quant_fusion backends base_static_graph caching collective_fusion compiler_interface counter cuda_graph decorators fix_functionalization fusion fusion_attn fx_utils inductor_pass matcher_utils monitor noop_elimination partition_rules pass_manager piecewise_backend post_cleanup qk_norm_rope_fusion rocm_aiter_fusion sequence_parallelism torch25_custom_graph_pass vllm_inductor_pass wrapper config config attention cache compilation device ec_transfer kv_events kv_transfer load lora model multimodal observability parallel pooler profiler scheduler speculative speech_to_text structured_outputs utils vllm device_allocator device_allocator cumem distributed distributed communication_op kv_events parallel_state tpu_distributed_utils utils device_communicators device_communicators all2all all_reduce_utils base_device_communicator cpu_communicator cuda_communicator cuda_wrapper custom_all_reduce mnnvl_compat pynccl pynccl_allocator pynccl_wrapper quick_all_reduce ray_communicator shm_broadcast shm_object_storage symm_mem tpu_communicator xpu_communicator ec_transfer ec_transfer ec_transfer_state ec_connector ec_connector base example_connector factory eplb eplb async_worker eplb_state rebalance_execute policy policy abstract default kv_transfer kv_transfer kv_transfer_state kv_connector kv_connector base factory utils v1 v1 base decode_bench_connector example_connector lmcache_connector lmcache_mp_connector metrics mooncake_connector multi_connector nixl_connector offloading_connector lmcache_integration lmcache_integration multi_process_adapter utils vllm_v1_adapter p2p p2p p2p_nccl_connector p2p_nccl_engine tensor_memory_pool engine engine arg_utils async_llm_engine llm_engine protocol entrypoints entrypoints api_server chat_utils constants context launcher llm logger renderer responses_utils score_utils ssl tool tool_server utils anthropic anthropic protocol serving_messages cli cli collect_env main openai run_batch serve types benchmark benchmark base latency main serve startup sweep throughput openai openai api_server cli_args orca_metrics protocol run_batch serving_chat serving_chat_stream_harmony serving_completion serving_engine serving_models serving_responses serving_transcription speech_to_text utils parser parser harmony_utils responses_parser pooling pooling classify classify api_router protocol serving embed embed api_router conftest protocol serving pooling pooling api_router protocol serving score score api_router protocol serving sagemaker sagemaker routes serve serve cache cache api_router disagg disagg api_router protocol serving elastic_ep elastic_ep api_router middleware instrumentator instrumentator health metrics server_info lora lora api_router profile profile api_router rlhf rlhf api_router rpc rpc api_router sleep sleep api_router tokenize tokenize api_router serving inputs inputs data parse preprocess logging_utils logging_utils dump_input formatter lazy log_time lora lora lora_model lora_weights model_manager peft_helper request resolver utils worker_manager layers layers base base_linear column_parallel_linear fused_moe logits_processor replicated_linear row_parallel_linear utils vocal_parallel_embedding ops ops ipex_ops ipex_ops lora_ops torch_ops torch_ops lora_ops triton_ops triton_ops fused_moe_lora_op kernel_utils lora_expand_op lora_kernel_metadata lora_shrink_op utils xla_ops xla_ops lora_ops punica_wrapper punica_wrapper punica_base punica_cpu punica_gpu punica_selector punica_tpu punica_xpu utils model_executor model_executor custom_op parameter utils layers layers activation attention_layer_base batch_invariant conv kda layernorm lightning_attn linear logits_processor mla pooler resampler utils vocab_parallel_embedding fla fla ops ops chunk chunk_delta_h chunk_o chunk_scaled_dot_kkt cumsum fused_recurrent index kda l2norm layernorm_guard op solve_tril utils wy_fast fused_moe fused_moe all2all_utils batched_deep_gemm_moe config cpu_fused_moe cutlass_moe deep_gemm_moe deep_gemm_utils deepep_ht_prepare_finalize deepep_ll_prepare_finalize flashinfer_cutedsl_moe flashinfer_cutlass_moe flashinfer_cutlass_prepare_finalize flashinfer_trtllm_moe fused_batched_moe fused_marlin_moe fused_moe fused_moe_method_base fused_moe_modular_method gpt_oss_triton_kernels_moe layer modular_kernel moe_align_block_size moe_pallas moe_permute_unpermute moe_torch_iterative pplx_prepare_finalize prepare_finalize rocm_aiter_fused_moe routing_simulator shared_fused_moe topk_weight_and_reduce triton_deep_gemm_moe trtllm_moe unquantized_fused_moe_method utils zero_expert_fused_moe mamba mamba abstract linear_attn mamba_mixer mamba_mixer2 mamba_utils short_conv ops ops causal_conv1d layernorm_gated mamba_ssm ssd_bmm ssd_chunk_scan ssd_chunk_state ssd_combined ssd_state_passing quantization quantization auto_round awq awq_marlin awq_triton base_config bitblas bitsandbytes cpu_wna16 deepspeedfp experts_int8 fbgemm_fp8 fp8 fp_quant gguf gptq gptq_bitblas gptq_marlin gptq_marlin_24 hqq_marlin inc input_quant_fp8 ipex_quant kv_cache modelopt moe_wna16 mxfp4 petit ptpc_fp8 qutlass_utils rtn schema torchao tpu_int8 compressed_tensors compressed_tensors compressed_tensors compressed_tensors_moe triton_scaled_mm utils schemes schemes compressed_tensors_24 compressed_tensors_scheme compressed_tensors_w4a4_nvfp4 compressed_tensors_w4a8_fp8 compressed_tensors_w4a8_int compressed_tensors_w4a16_24 compressed_tensors_w4a16_nvfp4 compressed_tensors_w8a8_fp8 compressed_tensors_w8a8_int8 compressed_tensors_w8a16_fp8 compressed_tensors_wNa16 transform transform linear module utils schemes schemes linear_qutlass_nvfp4 kernels kernels mixed_precision mixed_precision allspark bitblas conch cutlass dynamic_4bit exllama MPLinearKernel machete marlin xpu scaled_mm scaled_mm aiter cpu cutlass ScaledMMLinearKernel triton xla quark quark quark quark_moe utils schemes schemes quark_ocp_mx quark_scheme quark_w8a8_fp8 quark_w8a8_int8 utils utils allspark_utils bitblas_utils flashinfer_fp4_moe flashinfer_utils fp8_utils gptq_utils int8_utils layer_utils machete_utils marlin_utils marlin_utils_fp4 marlin_utils_fp8 marlin_utils_test marlin_utils_test_24 mxfp4_utils mxfp6_utils mxfp8_utils nvfp4_emulation_utils nvfp4_moe_support ocp_mx_utils petit_utils quant_utils w8a8_utils rotary_embedding rotary_embedding base common deepseek_scaling_rope dual_chunk_rope dynamic_ntk_alpha_rope dynamic_ntk_scaling_rope ernie45_vl_rope linear_scaling_rope llama3_rope llama4_vision_rope mrope ntk_scaling_rope phi3_long_rope_scaled_rope xdrope yarn_scaling_rope model_loader model_loader base_loader bitsandbytes_loader default_loader dummy_loader gguf_loader online_quantization runai_streamer_loader sharded_state_loader tensorizer tensorizer_loader tpu utils weight_utils models models adapters afmoe aimv2 apertus arcee arctic aria audioflamingo3 aya_vision bagel baichuan bailing_moe bamba bee bert bert_with_rope blip blip2 bloom chameleon chatglm clip cohere2_vision commandr config dbrx deepencoder deepseek_eagle deepseek_mtp deepseek_ocr deepseek_v2 deepseek_vl2 dots1 dots_ocr ernie45 ernie45_moe ernie45_vl ernie45_vl_moe ernie_mtp exaone exaone4 fairseq2_llama falcon falcon_h1 flex_olmo fuyu gemma gemma2 gemma3 gemma3_mm gemma3n gemma3n_mm glm glm4 glm4_1v glm4_moe glm4_moe_mtp glm4v gpt2 gpt_bigcode gpt_j gpt_neox gpt_oss granite granite_speech granitemoe granitemoehybrid granitemoeshared gritlm grok1 h2ovl hunyuan_v1 hunyuan_vision hyperclovax_vision idefics2_vision_model idefics3 interfaces interfaces_base intern_vit internlm2 internlm2_ve interns1 interns1_vit internvl jais jais2 jamba jina_vl keye keye_vl1_5 kimi_linear kimi_vl lfm2 lfm2_moe lightonocr llama llama4 llama4_eagle llama_eagle llama_eagle3 llava llava_next llava_next_video llava_onevision longcat_flash longcat_flash_mtp mamba mamba2 medusa midashenglm mimo mimo_mtp mimo_v2_flash minicpm minicpm3 minicpm_eagle minicpmo minicpmv minimax_m2 minimax_text_01 minimax_vl_01 mistral3 mistral_large_3 mistral_large_3_eagle mixtral mllama4 mlp_speculator modernbert module_mapping molmo moonvit mpt nano_nemotron_vl nemotron nemotron_h nemotron_nas nemotron_vl nvlm_d olmo olmo2 olmoe opencua openpangu openpangu_mtp opt orion ouro ovis ovis2_5 paddleocr_vl paligemma persimmon phi phi3 phi3v phi4mm phi4mm_audio phi4mm_utils phimoe pixtral plamo2 plamo3 qwen qwen2 qwen2_5_omni_thinker qwen2_5_vl qwen2_audio qwen2_moe qwen2_rm qwen2_vl qwen3 qwen3_moe qwen3_next qwen3_next_mtp qwen3_omni_moe_thinker qwen3_vl qwen3_vl_moe qwen_vl radio registry roberta rvl seed_oss siglip siglip2navit skyworkr1v smolvlm solar stablelm starcoder2 step3_text step3_vl swin tarsier telechat2 teleflm terratorch ultravox utils vision voxtral voxtral_streaming whisper whisper_utils zamba2 transformers transformers base causal legacy moe multimodal pooling utils warmup warmup deep_gemm_warmup kernel_warmup multimodal multimodal audio base cache evs hasher image inputs parse processing profiling registry utils video platforms platforms cpu cuda interface rocm tpu xpu plugins plugins io_processors io_processors interface lora_resolvers lora_resolvers filesystem_resolver profiler profiler layerwise_profile utils wrapper ray ray lazy_utils ray_env reasoning reasoning abs_reasoning_parsers basic_parsers deepseek_r1_reasoning_parser deepseek_v3_reasoning_parser ernie45_reasoning_parser glm4_moe_reasoning_parser gptoss_reasoning_parser granite_reasoning_parser holo2_reasoning_parser hunyuan_a13b_reasoning_parser identity_reasoning_parser minimax_m2_reasoning_parser mistral_reasoning_parser olmo3_reasoning_parser qwen3_reasoning_parser seedoss_reasoning_parser step3_reasoning_parser tokenizers tokenizers deepseek_v32 deepseek_v32_encoding detokenizer_utils hf mistral protocol registry tool_parsers tool_parsers abstract_tool_parser deepseekv3_tool_parser deepseekv31_tool_parser deepseekv32_tool_parser ernie45_tool_parser functiongemma_tool_parser gigachat3_tool_parser glm4_moe_tool_parser glm47_moe_tool_parser granite_20b_fc_tool_parser granite_tool_parser hermes_tool_parser hunyuan_a13b_tool_parser internlm2_tool_parser jamba_tool_parser kimi_k2_tool_parser llama4_pythonic_tool_parser llama_tool_parser longcat_tool_parser minimax_m2_tool_parser minimax_tool_parser mistral_tool_parser olmo3_tool_parser openai_tool_parser phi4mini_tool_parser pythonic_tool_parser qwen3coder_tool_parser qwen3xml_tool_parser seed_oss_tool_parser step3_tool_parser utils xlam_tool_parser transformers_utils transformers_utils config config_parser_base dynamic_module gguf_utils processor repo_utils runai_utils s3_utils tokenizer utils chat_templates chat_templates registry configs configs afmoe arctic bagel chatglm deepseek_vl2 dotsocr eagle falcon flex_olmo hunyuan_vl jais kimi_linear kimi_vl lfm2_moe medusa midashenglm mistral mlp_speculator moonvit nemotron nemotron_h olmo3 ovis qwen3_next radio step3_vl tarsier2 ultravox speculators speculators algos base processors processors bagel deepseek_ocr deepseek_vl2 hunyuan_vl hunyuan_vl_image ovis ovis2_5 triton_utils triton_utils importing usage usage usage_lib utils utils argparse_utils async_utils cache collection_utils counter deep_gemm flashinfer func_utils gc_utils hashing import_utils jsontree math_utils mem_constants mem_utils nccl network_utils nvtx_pytorch_hooks platform_utils profiling registry serial_utils system_utils tensor_schema torch_utils v1 v1 cudagraph_dispatcher kv_cache_interface outputs request serial_utils utils attention attention backends backends cpu_attn flash_attn flashinfer flex_attention gdn_attn linear_attn mamba1_attn mamba2_attn mamba_attn pallas rocm_aiter_fa rocm_aiter_unified_attn rocm_attn short_conv_attn tree_attn triton_attn utils mla mla aiter_triton_mla common cutlass_mla flashattn_mla flashinfer_mla flashmla flashmla_sparse indexer rocm_aiter_mla rocm_aiter_mla_sparse triton_mla core core block_pool encoder_cache_manager kv_cache_coordinator kv_cache_manager kv_cache_metrics kv_cache_utils single_type_kv_cache_manager sched sched async_scheduler interface output request_queue scheduler utils engine engine async_llm coordinator core core_client detokenizer exceptions input_processor llm_engine logprobs output_processor parallel_sampling utils executor executor abstract multiproc_executor ray_distributed_executor ray_executor ray_utils uniproc_executor kv_offload kv_offload abstract arc_manager backend cpu factory lru_manager mediums spec backends backends cpu worker worker cpu_gpu worker metrics metrics loggers perf prometheus ray_wrappers reader stats pool pool metadata sample sample metadata rejection_sampler sampler logits_processor logits_processor builtin interface state ops ops bad_words logprobs penalties topk_topp_sampler tpu tpu metadata sampler spec_decode spec_decode eagle medusa metadata metrics ngram_proposer suffix_decoding utils structured_output structured_output backend_guidance backend_lm_format_enforcer backend_outlines backend_types backend_xgrammar request utils worker worker block_table cp_utils cpu_model_runner cpu_worker dp_utils ec_connector_model_runner_mixin gpu_input_batch gpu_model_runner gpu_ubatch_wrapper gpu_worker kv_connector_model_runner_mixin lora_model_runner_mixin tpu_input_batch tpu_model_runner tpu_worker ubatch_utils ubatching utils worker_base workspace xpu_model_runner xpu_worker gpu gpu async_utils attn_utils block_table cudagraph_utils dp_utils input_batch model_runner states structured_outputs metrics metrics logits sample sample gumbel logprob metadata min_p output penalties sampler spec_decode spec_decode eagle eagle_cudagraph rejection_sample CLI Reference CLI Reference vllm serve vllm chat vllm complete vllm run-batch vllm bench vllm bench vllm bench latency vllm bench serve vllm bench sweep plot vllm bench sweep plot_pareto vllm bench sweep serve vllm bench sweep serve_sla vllm bench throughput Community Community Contact Us Meetups Sponsors Governance Governance Collaboration Policy Committers Governance Process Blog Forum Slack Table of contents Quark Installation Quantization Process 1. Load the Model 2. Prepare the Calibration Dataloader 3. Set the Quantization Configuration 4. Quantize the Model and Export 5. Evaluation in vLLM Quark Quantization Script Using OCP MX (MXFP4, MXFP6) models Using Quark Quantized layerwise Auto Mixed Precision (AMP) Models 1. Quantize a model using mixed precision in AMD Quark 2. inference the quantized mixed precision model in vLLM AMD Quark¶ Quantization can effectively reduce memory and bandwidth usage, accelerate computation and improve throughput while with minimal accuracy loss. vLLM can leverage Quark, the flexible and powerful quantization toolkit, to produce performant quantized models to run on AMD GPUs. Quark has specialized support for quantizing large language models with weight, activation and kv-cache quantization and cutting-edge quantization algorithms like AWQ, GPTQ, Rotation and SmoothQuant. Quark Installation¶ Before quantizing models, you need to install Quark. The latest release of Quark can be installed with pip: pip install amd-quark You can refer to Quark installation guide for more installation details. Additionally, install vllm and lm-evaluation-harness for evaluation: pip install vllm "lm-eval[api]>=0.4.9.2" Quantization Process¶ After installing Quark, we will use an example to illustrate how to use Quark. The Quark quantization process can be listed for 5 steps as below: Load the model Prepare the calibration dataloader Set the quantization configuration Quantize the model and export Evaluation in vLLM 1. Load the Model¶ Quark uses Transformers to fetch model and tokenizer. Code from transformers import AutoTokenizer, AutoModelForCausalLM MODEL_ID = "meta-llama/Llama-2-70b-chat-hf" MAX_SEQ_LEN = 512 model = AutoModelForCausalLM.from_pretrained( MODEL_ID, device_map="auto", dtype="auto", ) model.eval() tokenizer = AutoTokenizer.from_pretrained(MODEL_ID, model_max_length=MAX_SEQ_LEN) tokenizer.pad_token = tokenizer.eos_token 2. Prepare the Calibration Dataloader¶ Quark uses the PyTorch Dataloader to load calibration data. For more details about how to use calibration datasets efficiently, please refer to Adding Calibration Datasets. Code from datasets import load_dataset from torch.utils.data import DataLoader BATCH_SIZE = 1 NUM_CALIBRATION_DATA = 512 # Load the dataset and get calibration data. dataset = load_dataset("mit-han-lab/pile-val-backup", split="validation") text_data = dataset["text"][:NUM_CALIBRATION_DATA] tokenized_outputs = tokenizer( text_data, return_tensors="pt", padding=True, truncation=True, max_length=MAX_SEQ_LEN, ) calib_dataloader = DataLoader( tokenized_outputs['input_ids'], batch_size=BATCH_SIZE, drop_last=True, ) 3. Set the Quantization Configuration¶ We need to set the quantization configuration, you can check quark config guide for further details. Here we use FP8 per-tensor quantization on weight, activation, kv-cache and the quantization algorithm is AutoSmoothQuant. Note Note the quantization algorithm needs a JSON config file and the config file is located in Quark Pytorch examples, under the directory examples/torch/language_modeling/llm_ptq/models. For example, AutoSmoothQuant config file for Llama is examples/torch/language_modeling/llm_ptq/models/llama/autosmoothquant_config.json. Code from quark.torch.quantization import (Config, QuantizationConfig, FP8E4M3PerTensorSpec, load_quant_algo_config_from_file) # Define fp8/per-tensor/static spec. FP8_PER_TENSOR_SPEC = FP8E4M3PerTensorSpec( observer_method="min_max", is_dynamic=False, ).to_quantization_spec() # Define global quantization config, input tensors and weight apply FP8_PER_TENSOR_SPEC. global_quant_config = QuantizationConfig( input_tensors=FP8_PER_TENSOR_SPEC, weight=FP8_PER_TENSOR_SPEC, ) # Define quantization config for kv-cache layers, output tensors apply FP8_PER_TENSOR_SPEC. KV_CACHE_SPEC = FP8_PER_TENSOR_SPEC kv_cache_layer_names_for_llama = ["*k_proj", "*v_proj"] kv_cache_quant_config = { name: QuantizationConfig( input_tensors=global_quant_config.input_tensors, weight=global_quant_config.weight, output_tensors=KV_CACHE_SPEC, ) for name in kv_cache_layer_names_for_llama } layer_quant_config = kv_cache_quant_config.copy() # Define algorithm config by config file. LLAMA_AUTOSMOOTHQUANT_CONFIG_FILE = "examples/torch/language_modeling/llm_ptq/models/llama/autosmoothquant_config.json" algo_config = load_quant_algo_config_from_file(LLAMA_AUTOSMOOTHQUANT_CONFIG_FILE) EXCLUDE_LAYERS = ["lm_head"] quant_config = Config( global_quant_config=global_quant_config, layer_quant_config=layer_quant_config, kv_cache_quant_config=kv_cache_quant_config, exclude=EXCLUDE_LAYERS, algo_config=algo_config, ) 4. Quantize the Model and Export¶ Then we can apply the quantization. After quantizing, we need to freeze the quantized model first before exporting. Note that we need to export model with format of HuggingFace safetensors, you can refer to HuggingFace format exporting for more exporting format details. Code import torch from quark.torch import ModelQuantizer, ModelExporter from quark.torch.export import ExporterConfig, JsonExporterConfig # Apply quantization. quantizer = ModelQuantizer(quant_config) quant_model = quantizer.quantize_model(model, calib_dataloader) # Freeze quantized model to export. freezed_model = quantizer.freeze(model) # Define export config. LLAMA_KV_CACHE_GROUP = ["*k_proj", "*v_proj"] export_config = ExporterConfig(json_export_config=JsonExporterConfig()) export_config.json_export_config.kv_cache_group = LLAMA_KV_CACHE_GROUP # Model: Llama-2-70b-chat-hf-w-fp8-a-fp8-kvcache-fp8-pertensor-autosmoothquant EXPORT_DIR = MODEL_ID.split("/")[1] + "-w-fp8-a-fp8-kvcache-fp8-pertensor-autosmoothquant" exporter = ModelExporter(config=export_config, export_dir=EXPORT_DIR) with torch.no_grad(): exporter.export_safetensors_model( freezed_model, quant_config=quant_config, tokenizer=tokenizer, ) 5. Evaluation in vLLM¶ Now, you can load and run the Quark quantized model directly through the LLM entrypoint: Code from vllm import LLM, SamplingParams # Sample prompts. prompts = [ "Hello, my name is", "The president of the United States is", "The capital of France is", "The future of AI is", ] # Create a sampling params object. sampling_params = SamplingParams(temperature=0.8, top_p=0.95) # Create an LLM. llm = LLM( model="Llama-2-70b-chat-hf-w-fp8-a-fp8-kvcache-fp8-pertensor-autosmoothquant", kv_cache_dtype="fp8", quantization="quark", ) # Generate texts from the prompts. The output is a list of RequestOutput objects # that contain the prompt, generated text, and other information. outputs = llm.generate(prompts, sampling_params) # Print the outputs. print("\nGenerated Outputs:\n" + "-" * 60) for output in outputs: prompt = output.prompt generated_text = output.outputs[0].text print(f"Prompt: {prompt!r}") print(f"Output: {generated_text!r}") print("-" * 60) Or, you can use lm_eval to evaluate accuracy: lm_eval --model vllm \ --model_args pretrained=Llama-2-70b-chat-hf-w-fp8-a-fp8-kvcache-fp8-pertensor-autosmoothquant,kv_cache_dtype='fp8',quantization='quark' \ --tasks gsm8k Quark Quantization Script¶ In addition to the example of Python API above, Quark also offers a quantization script to quantize large language models more conveniently. It supports quantizing models with variety of different quantization schemes and optimization algorithms. It can export the quantized model and run evaluation tasks on the fly. With the script, the example above can be: python3 quantize_quark.py --model_dir meta-llama/Llama-2-70b-chat-hf \ --output_dir /path/to/output \ --quant_scheme w_fp8_a_fp8 \ --kv_cache_dtype fp8 \ --quant_algo autosmoothquant \ --num_calib_data 512 \ --model_export hf_format \ --tasks gsm8k Using OCP MX (MXFP4, MXFP6) models¶ vLLM supports loading MXFP4 and MXFP6 models quantized offline through AMD Quark, compliant with Open Compute Project (OCP) specification. The scheme currently only supports dynamic quantization for activations. Example usage, after installing the latest AMD Quark release: vllm serve fxmarty/qwen_1.5-moe-a2.7b-mxfp4 --tensor-parallel-size 1 # or, for a model using fp6 activations and fp4 weights: vllm serve fxmarty/qwen1.5_moe_a2.7b_chat_w_fp4_a_fp6_e2m3 --tensor-parallel-size 1 A simulation of the matrix multiplication execution in MXFP4/MXFP6 can be run on devices that do not support OCP MX operations natively (e.g. AMD Instinct MI325, MI300 and MI250), dequantizing weights from FP4/FP6 to half precision on the fly, using a fused kernel. This is useful e.g. to evaluate FP4/FP6 models using vLLM, or alternatively to benefit from the ~2.5-4x memory savings (compared to float16 and bfloat16). To generate offline models quantized using MXFP4 data type, the easiest approach is to use AMD Quark's quantization script, as an example: python quantize_quark.py --model_dir Qwen/Qwen1.5-MoE-A2.7B-Chat \ --quant_scheme w_mxfp4_a_mxfp4 \ --output_dir qwen_1.5-moe-a2.7b-mxfp4 \ --skip_evaluation \ --model_export hf_format \ --group_size 32 The current integration supports all combination of FP4, FP6_E3M2, FP6_E2M3 used for either weights or activations. Using Quark Quantized layerwise Auto Mixed Precision (AMP) Models¶ vLLM also supports loading layerwise mixed precision model quantized using AMD Quark. Currently, mixed scheme of {MXFP4, FP8} is supported, where FP8 here denotes for FP8 per-tensor scheme. More mixed precision schemes are planned to be supported in a near future, including Unquantized Linear and/or MoE layer(s) as an option for each layer, i.e., mixed of {MXFP4, FP8, BF16/FP16} MXFP6 quantization extension, i.e., {MXFP4, MXFP6, FP8, BF16/FP16} Although one can maximize serving throughput using the lowest precision supported on a given device (e.g. MXFP4 for AMD Instinct MI355, FP8 for AMD Instinct MI300), these aggressive schemes can be detrimental to accuracy recovering from quantization on target tasks. Mixed precision allows to strike a balance between maximizing accuracy and throughput. There are two steps to generate and deploy a mixed precision model quantized with AMD Quark, as shown below. 1. Quantize a model using mixed precision in AMD Quark¶ Firstly, the layerwise mixed-precision configuration for a given LLM model is searched and then quantized using AMD Quark. We will provide a detailed tutorial with Quark APIs later. As examples, we provide some ready-to-use quantized mixed precision model to show the usage in vLLM and the accuracy benefits. They are: amd/Llama-2-70b-chat-hf-WMXFP4FP8-AMXFP4FP8-AMP-KVFP8 amd/Mixtral-8x7B-Instruct-v0.1-WMXFP4FP8-AMXFP4FP8-AMP-KVFP8 amd/Qwen3-8B-WMXFP4FP8-AMXFP4FP8-AMP-KVFP8 2. inference the quantized mixed precision model in vLLM¶ Models quantized with AMD Quark using mixed precision can natively be reload in vLLM, and e.g. evaluated using lm-evaluation-harness as follows: lm_eval --model vllm \ --model_args pretrained=amd/Llama-2-70b-chat-hf-WMXFP4FP8-AMXFP4FP8-AMP-KVFP8,tensor_parallel_size=4,dtype=auto,gpu_memory_utilization=0.8,trust_remote_code=False \ --tasks mmlu \ --batch_size auto December 24, 2025 ``` pip install amd-quark ``` **Pattern 2:** To generate offline models quantized using MXFP4 data type, the easiest approach is to use AMD Quark's quantization script, as an example: ``` python quantize_quark.py --model_dir Qwen/Qwen1.5-MoE-A2.7B-Chat \ --quant_scheme w_mxfp4_a_mxfp4 \ --output_dir qwen_1.5-moe-a2.7b-mxfp4 \ --skip_evaluation \ --model_export hf_format \ --group_size 32 ``` **Pattern 3:** vLLM GitHub Home User Guide User Guide Getting Started Getting Started Quickstart Installation Installation GPU CPU TPU Examples Examples Offline inference Offline inference Async LLM Streaming Audio Language Automatic Prefix Caching Basic Batch LLM Inference Chat With Tools Context Extension Data Parallel Disaggregated Prefill V1 Disaggregated Prefill Encoder Decoder Multimodal KV Load Failure Recovery Test LLM Engine Example LLM Engine Reset Kv Load Sharded State Logits Processor LoRA With Quantization Inference Metrics Mistral-Small MLPSpeculator MultiLoRA Inference Offline Inference with the OpenAI Batch file format Prefix Caching Prompt Embed Inference Qwen2.5-Omni Offline Inference Examples Qwen3 Omni Qwen 1M Reproducibility RLHF RLHF Colocate RLHF Online Quant RLHF Utils Save Sharded State Simple Profiling Skip Loading Weights In Engine Init Spec Decode Structured Outputs Torchrun Dp Example Torchrun Example Vision Language Vision Language Multi Image Online serving Online serving API Client Helm Charts Monitoring Dashboards Disaggregated Encoder Disaggregated Prefill Disaggregated Serving Disaggregated Serving P2P Nccl Xpyd Elastic Ep Gradio OpenAI Chatbot Webserver Gradio Webserver Kv Events Subscriber Multi-Node-Serving Multi Instance Data Parallel OpenAI Chat Completion Client OpenAI Chat Completion Client For Multimodal OpenAI Chat Completion Client With Tools OpenAI Chat Completion Client With Tools Required OpenAI Chat Completion Client With Tools Xlam OpenAI Chat Completion Client With Tools Xlam Streaming OpenAI Chat Completion Tool Calls With Reasoning OpenAI Chat Completion With Reasoning OpenAI Chat Completion With Reasoning Streaming OpenAI Completion Client OpenAI Responses Client OpenAI Responses Client With Mcp Tools OpenAI Responses Client With Tools OpenAI Transcription Client OpenAI Translation Client Setup OpenTelemetry POC Prometheus and Grafana Prompt Embed Inference With OpenAI Client Ray Serve Deepseek Retrieval Augmented Generation With Langchain Retrieval Augmented Generation With Llamaindex Run Cluster Sagemaker-Entrypoint Streamlit OpenAI Chatbot Webserver Structured Outputs Token Generation Client Utils Others Others LMCache Examples Logging Configuration Tensorize vLLM Model Pooling Pooling Classify Embed Plugin Pooling Score Token Classify Token Embed General General vLLM V1 Frequently Asked Questions Production Metrics Reproducibility Security Troubleshooting Usage Stats Collection Inference and Serving Inference and Serving Offline Inference OpenAI-Compatible Server Context Parallel Deployment Data Parallel Deployment Troubleshooting distributed deployments Expert Parallel Deployment Parallelism and Scaling Integrations Integrations LangChain LlamaIndex Deployment Deployment Using Docker Using Kubernetes Using Nginx Frameworks Frameworks Anyscale AnythingLLM AnythingLLM Table of contents Prerequisites Deploy AutoGen BentoML Cerebrium Chatbox Dify dstack Haystack Helm Hugging Face Inference Endpoints LiteLLM Lobe Chat LWS Modal Open WebUI Retrieval-Augmented Generation SkyPilot Streamlit NVIDIA Triton Integrations Integrations KAITO KServe Kthena KubeAI KubeRay Llama Stack llm-d llmaz Production stack Training Training Reinforcement Learning from Human Feedback Transformers Reinforcement Learning Configuration Configuration Conserving Memory Engine Arguments Environment Variables Model Resolution Optimization and Tuning Server Arguments TPU Models Models Supported Models Generative Models Pooling Models Extensions Extensions Loading Model weights with fastsafetensors Loading models with Run:ai Model Streamer Loading models with CoreWeave's Tensorizer Hardware Supported Models Hardware Supported Models CPU - Intel® Xeon® XPU - Intel® GPUs TPU Features Features Automatic Prefix Caching Batch Invariance Custom Arguments Custom Logits Processors Disaggregated Encoder Disaggregated Prefilling (experimental) Interleaved Thinking LoRA Adapters MooncakeConnector Usage Guide Multimodal Inputs NixlConnector Usage Guide Prompt Embedding Inputs Reasoning Outputs Sleep Mode Speculative Decoding Structured Outputs Tool Calling Quantization Quantization AutoAWQ AutoRound BitBLAS BitsAndBytes FP8 W8A8 GGUF GPTQModel FP8 INC INT4 W4A16 INT8 W8A8 NVIDIA Model Optimizer Quantized KV Cache AMD Quark TorchAO Developer Guide Developer Guide General General Deprecation Policy Dockerfile Incremental Compilation Workflow Profiling vLLM Vulnerability Management Model Implementation Model Implementation Basic Model Registering a Model Unit Testing Multi-Modal Support Speech-to-Text (Transcription/Translation) Support CI CI CI Failures Nightly Builds of vLLM Wheels Update PyTorch version on vLLM OSS CI/CD Design Documents Design Documents Plugins Plugins IO Processor Plugins LoRA Resolver Plugins Plugin System Architecture Overview CUDA Graphs Dual Batch Overlap How to debug the vLLM-torch.compile integration Fused MoE Modular Kernel Integration with Hugging Face Hybrid KV Cache Manager Logits Processors Metrics Multi-Modal Data Processing Fused MoE Kernel Features Python Multiprocessing Optimization levels P2P NCCL Connector Paged Attention Automatic Prefix Caching torch.compile integration Benchmarking Benchmarking Benchmark CLI Parameter Sweeps Performance Dashboard API Reference API Reference vllm vllm beam_search collect_env connections env_override envs forward_context logger logits_process logprobs outputs pooling_params sampling_params scalar_type scripts sequence tasks tracing version assets assets audio base image video attention attention layer selector backends backends abstract registry utils layers layers chunked_local_attention cross_attention encoder_only_attention mm_encoder_attention ops ops chunked_prefill_paged_decode common flashmla merge_attn_states paged_attn pallas_kv_cache_update prefix_prefill rocm_aiter_mla_sparse triton_decode_attention triton_merge_attn_states triton_reshape_and_cache_flash triton_unified_attention vit_attn_wrappers utils utils fa_utils kv_sharing_utils kv_transfer_utils benchmarks benchmarks datasets latency serve startup throughput lib lib endpoint_request_func ready_checker utils sweep sweep cli param_sweep plot plot_pareto serve serve_sla server sla_sweep utils compilation compilation activation_quant_fusion backends base_static_graph caching collective_fusion compiler_interface counter cuda_graph decorators fix_functionalization fusion fusion_attn fx_utils inductor_pass matcher_utils monitor noop_elimination partition_rules pass_manager piecewise_backend post_cleanup qk_norm_rope_fusion rocm_aiter_fusion sequence_parallelism torch25_custom_graph_pass vllm_inductor_pass wrapper config config attention cache compilation device ec_transfer kv_events kv_transfer load lora model multimodal observability parallel pooler profiler scheduler speculative speech_to_text structured_outputs utils vllm device_allocator device_allocator cumem distributed distributed communication_op kv_events parallel_state tpu_distributed_utils utils device_communicators device_communicators all2all all_reduce_utils base_device_communicator cpu_communicator cuda_communicator cuda_wrapper custom_all_reduce mnnvl_compat pynccl pynccl_allocator pynccl_wrapper quick_all_reduce ray_communicator shm_broadcast shm_object_storage symm_mem tpu_communicator xpu_communicator ec_transfer ec_transfer ec_transfer_state ec_connector ec_connector base example_connector factory eplb eplb async_worker eplb_state rebalance_execute policy policy abstract default kv_transfer kv_transfer kv_transfer_state kv_connector kv_connector base factory utils v1 v1 base decode_bench_connector example_connector lmcache_connector lmcache_mp_connector metrics mooncake_connector multi_connector nixl_connector offloading_connector lmcache_integration lmcache_integration multi_process_adapter utils vllm_v1_adapter p2p p2p p2p_nccl_connector p2p_nccl_engine tensor_memory_pool engine engine arg_utils async_llm_engine llm_engine protocol entrypoints entrypoints api_server chat_utils constants context launcher llm logger renderer responses_utils score_utils ssl tool tool_server utils anthropic anthropic protocol serving_messages cli cli collect_env main openai run_batch serve types benchmark benchmark base latency main serve startup sweep throughput openai openai api_server cli_args orca_metrics protocol run_batch serving_chat serving_chat_stream_harmony serving_completion serving_engine serving_models serving_responses serving_transcription speech_to_text utils parser parser harmony_utils responses_parser pooling pooling classify classify api_router protocol serving embed embed api_router conftest protocol serving pooling pooling api_router protocol serving score score api_router protocol serving sagemaker sagemaker routes serve serve cache cache api_router disagg disagg api_router protocol serving elastic_ep elastic_ep api_router middleware instrumentator instrumentator health metrics server_info lora lora api_router profile profile api_router rlhf rlhf api_router rpc rpc api_router sleep sleep api_router tokenize tokenize api_router serving inputs inputs data parse preprocess logging_utils logging_utils dump_input formatter lazy log_time lora lora lora_model lora_weights model_manager peft_helper request resolver utils worker_manager layers layers base base_linear column_parallel_linear fused_moe logits_processor replicated_linear row_parallel_linear utils vocal_parallel_embedding ops ops ipex_ops ipex_ops lora_ops torch_ops torch_ops lora_ops triton_ops triton_ops fused_moe_lora_op kernel_utils lora_expand_op lora_kernel_metadata lora_shrink_op utils xla_ops xla_ops lora_ops punica_wrapper punica_wrapper punica_base punica_cpu punica_gpu punica_selector punica_tpu punica_xpu utils model_executor model_executor custom_op parameter utils layers layers activation attention_layer_base batch_invariant conv kda layernorm lightning_attn linear logits_processor mla pooler resampler utils vocab_parallel_embedding fla fla ops ops chunk chunk_delta_h chunk_o chunk_scaled_dot_kkt cumsum fused_recurrent index kda l2norm layernorm_guard op solve_tril utils wy_fast fused_moe fused_moe all2all_utils batched_deep_gemm_moe config cpu_fused_moe cutlass_moe deep_gemm_moe deep_gemm_utils deepep_ht_prepare_finalize deepep_ll_prepare_finalize flashinfer_cutedsl_moe flashinfer_cutlass_moe flashinfer_cutlass_prepare_finalize flashinfer_trtllm_moe fused_batched_moe fused_marlin_moe fused_moe fused_moe_method_base fused_moe_modular_method gpt_oss_triton_kernels_moe layer modular_kernel moe_align_block_size moe_pallas moe_permute_unpermute moe_torch_iterative pplx_prepare_finalize prepare_finalize rocm_aiter_fused_moe routing_simulator shared_fused_moe topk_weight_and_reduce triton_deep_gemm_moe trtllm_moe unquantized_fused_moe_method utils zero_expert_fused_moe mamba mamba abstract linear_attn mamba_mixer mamba_mixer2 mamba_utils short_conv ops ops causal_conv1d layernorm_gated mamba_ssm ssd_bmm ssd_chunk_scan ssd_chunk_state ssd_combined ssd_state_passing quantization quantization auto_round awq awq_marlin awq_triton base_config bitblas bitsandbytes cpu_wna16 deepspeedfp experts_int8 fbgemm_fp8 fp8 fp_quant gguf gptq gptq_bitblas gptq_marlin gptq_marlin_24 hqq_marlin inc input_quant_fp8 ipex_quant kv_cache modelopt moe_wna16 mxfp4 petit ptpc_fp8 qutlass_utils rtn schema torchao tpu_int8 compressed_tensors compressed_tensors compressed_tensors compressed_tensors_moe triton_scaled_mm utils schemes schemes compressed_tensors_24 compressed_tensors_scheme compressed_tensors_w4a4_nvfp4 compressed_tensors_w4a8_fp8 compressed_tensors_w4a8_int compressed_tensors_w4a16_24 compressed_tensors_w4a16_nvfp4 compressed_tensors_w8a8_fp8 compressed_tensors_w8a8_int8 compressed_tensors_w8a16_fp8 compressed_tensors_wNa16 transform transform linear module utils schemes schemes linear_qutlass_nvfp4 kernels kernels mixed_precision mixed_precision allspark bitblas conch cutlass dynamic_4bit exllama MPLinearKernel machete marlin xpu scaled_mm scaled_mm aiter cpu cutlass ScaledMMLinearKernel triton xla quark quark quark quark_moe utils schemes schemes quark_ocp_mx quark_scheme quark_w8a8_fp8 quark_w8a8_int8 utils utils allspark_utils bitblas_utils flashinfer_fp4_moe flashinfer_utils fp8_utils gptq_utils int8_utils layer_utils machete_utils marlin_utils marlin_utils_fp4 marlin_utils_fp8 marlin_utils_test marlin_utils_test_24 mxfp4_utils mxfp6_utils mxfp8_utils nvfp4_emulation_utils nvfp4_moe_support ocp_mx_utils petit_utils quant_utils w8a8_utils rotary_embedding rotary_embedding base common deepseek_scaling_rope dual_chunk_rope dynamic_ntk_alpha_rope dynamic_ntk_scaling_rope ernie45_vl_rope linear_scaling_rope llama3_rope llama4_vision_rope mrope ntk_scaling_rope phi3_long_rope_scaled_rope xdrope yarn_scaling_rope model_loader model_loader base_loader bitsandbytes_loader default_loader dummy_loader gguf_loader online_quantization runai_streamer_loader sharded_state_loader tensorizer tensorizer_loader tpu utils weight_utils models models adapters afmoe aimv2 apertus arcee arctic aria audioflamingo3 aya_vision bagel baichuan bailing_moe bamba bee bert bert_with_rope blip blip2 bloom chameleon chatglm clip cohere2_vision commandr config dbrx deepencoder deepseek_eagle deepseek_mtp deepseek_ocr deepseek_v2 deepseek_vl2 dots1 dots_ocr ernie45 ernie45_moe ernie45_vl ernie45_vl_moe ernie_mtp exaone exaone4 fairseq2_llama falcon falcon_h1 flex_olmo fuyu gemma gemma2 gemma3 gemma3_mm gemma3n gemma3n_mm glm glm4 glm4_1v glm4_moe glm4_moe_mtp glm4v gpt2 gpt_bigcode gpt_j gpt_neox gpt_oss granite granite_speech granitemoe granitemoehybrid granitemoeshared gritlm grok1 h2ovl hunyuan_v1 hunyuan_vision hyperclovax_vision idefics2_vision_model idefics3 interfaces interfaces_base intern_vit internlm2 internlm2_ve interns1 interns1_vit internvl jais jais2 jamba jina_vl keye keye_vl1_5 kimi_linear kimi_vl lfm2 lfm2_moe lightonocr llama llama4 llama4_eagle llama_eagle llama_eagle3 llava llava_next llava_next_video llava_onevision longcat_flash longcat_flash_mtp mamba mamba2 medusa midashenglm mimo mimo_mtp mimo_v2_flash minicpm minicpm3 minicpm_eagle minicpmo minicpmv minimax_m2 minimax_text_01 minimax_vl_01 mistral3 mistral_large_3 mistral_large_3_eagle mixtral mllama4 mlp_speculator modernbert module_mapping molmo moonvit mpt nano_nemotron_vl nemotron nemotron_h nemotron_nas nemotron_vl nvlm_d olmo olmo2 olmoe opencua openpangu openpangu_mtp opt orion ouro ovis ovis2_5 paddleocr_vl paligemma persimmon phi phi3 phi3v phi4mm phi4mm_audio phi4mm_utils phimoe pixtral plamo2 plamo3 qwen qwen2 qwen2_5_omni_thinker qwen2_5_vl qwen2_audio qwen2_moe qwen2_rm qwen2_vl qwen3 qwen3_moe qwen3_next qwen3_next_mtp qwen3_omni_moe_thinker qwen3_vl qwen3_vl_moe qwen_vl radio registry roberta rvl seed_oss siglip siglip2navit skyworkr1v smolvlm solar stablelm starcoder2 step3_text step3_vl swin tarsier telechat2 teleflm terratorch ultravox utils vision voxtral voxtral_streaming whisper whisper_utils zamba2 transformers transformers base causal legacy moe multimodal pooling utils warmup warmup deep_gemm_warmup kernel_warmup multimodal multimodal audio base cache evs hasher image inputs parse processing profiling registry utils video platforms platforms cpu cuda interface rocm tpu xpu plugins plugins io_processors io_processors interface lora_resolvers lora_resolvers filesystem_resolver profiler profiler layerwise_profile utils wrapper ray ray lazy_utils ray_env reasoning reasoning abs_reasoning_parsers basic_parsers deepseek_r1_reasoning_parser deepseek_v3_reasoning_parser ernie45_reasoning_parser glm4_moe_reasoning_parser gptoss_reasoning_parser granite_reasoning_parser holo2_reasoning_parser hunyuan_a13b_reasoning_parser identity_reasoning_parser minimax_m2_reasoning_parser mistral_reasoning_parser olmo3_reasoning_parser qwen3_reasoning_parser seedoss_reasoning_parser step3_reasoning_parser tokenizers tokenizers deepseek_v32 deepseek_v32_encoding detokenizer_utils hf mistral protocol registry tool_parsers tool_parsers abstract_tool_parser deepseekv3_tool_parser deepseekv31_tool_parser deepseekv32_tool_parser ernie45_tool_parser functiongemma_tool_parser gigachat3_tool_parser glm4_moe_tool_parser glm47_moe_tool_parser granite_20b_fc_tool_parser granite_tool_parser hermes_tool_parser hunyuan_a13b_tool_parser internlm2_tool_parser jamba_tool_parser kimi_k2_tool_parser llama4_pythonic_tool_parser llama_tool_parser longcat_tool_parser minimax_m2_tool_parser minimax_tool_parser mistral_tool_parser olmo3_tool_parser openai_tool_parser phi4mini_tool_parser pythonic_tool_parser qwen3coder_tool_parser qwen3xml_tool_parser seed_oss_tool_parser step3_tool_parser utils xlam_tool_parser transformers_utils transformers_utils config config_parser_base dynamic_module gguf_utils processor repo_utils runai_utils s3_utils tokenizer utils chat_templates chat_templates registry configs configs afmoe arctic bagel chatglm deepseek_vl2 dotsocr eagle falcon flex_olmo hunyuan_vl jais kimi_linear kimi_vl lfm2_moe medusa midashenglm mistral mlp_speculator moonvit nemotron nemotron_h olmo3 ovis qwen3_next radio step3_vl tarsier2 ultravox speculators speculators algos base processors processors bagel deepseek_ocr deepseek_vl2 hunyuan_vl hunyuan_vl_image ovis ovis2_5 triton_utils triton_utils importing usage usage usage_lib utils utils argparse_utils async_utils cache collection_utils counter deep_gemm flashinfer func_utils gc_utils hashing import_utils jsontree math_utils mem_constants mem_utils nccl network_utils nvtx_pytorch_hooks platform_utils profiling registry serial_utils system_utils tensor_schema torch_utils v1 v1 cudagraph_dispatcher kv_cache_interface outputs request serial_utils utils attention attention backends backends cpu_attn flash_attn flashinfer flex_attention gdn_attn linear_attn mamba1_attn mamba2_attn mamba_attn pallas rocm_aiter_fa rocm_aiter_unified_attn rocm_attn short_conv_attn tree_attn triton_attn utils mla mla aiter_triton_mla common cutlass_mla flashattn_mla flashinfer_mla flashmla flashmla_sparse indexer rocm_aiter_mla rocm_aiter_mla_sparse triton_mla core core block_pool encoder_cache_manager kv_cache_coordinator kv_cache_manager kv_cache_metrics kv_cache_utils single_type_kv_cache_manager sched sched async_scheduler interface output request_queue scheduler utils engine engine async_llm coordinator core core_client detokenizer exceptions input_processor llm_engine logprobs output_processor parallel_sampling utils executor executor abstract multiproc_executor ray_distributed_executor ray_executor ray_utils uniproc_executor kv_offload kv_offload abstract arc_manager backend cpu factory lru_manager mediums spec backends backends cpu worker worker cpu_gpu worker metrics metrics loggers perf prometheus ray_wrappers reader stats pool pool metadata sample sample metadata rejection_sampler sampler logits_processor logits_processor builtin interface state ops ops bad_words logprobs penalties topk_topp_sampler tpu tpu metadata sampler spec_decode spec_decode eagle medusa metadata metrics ngram_proposer suffix_decoding utils structured_output structured_output backend_guidance backend_lm_format_enforcer backend_outlines backend_types backend_xgrammar request utils worker worker block_table cp_utils cpu_model_runner cpu_worker dp_utils ec_connector_model_runner_mixin gpu_input_batch gpu_model_runner gpu_ubatch_wrapper gpu_worker kv_connector_model_runner_mixin lora_model_runner_mixin tpu_input_batch tpu_model_runner tpu_worker ubatch_utils ubatching utils worker_base workspace xpu_model_runner xpu_worker gpu gpu async_utils attn_utils block_table cudagraph_utils dp_utils input_batch model_runner states structured_outputs metrics metrics logits sample sample gumbel logprob metadata min_p output penalties sampler spec_decode spec_decode eagle eagle_cudagraph rejection_sample CLI Reference CLI Reference vllm serve vllm chat vllm complete vllm run-batch vllm bench vllm bench vllm bench latency vllm bench serve vllm bench sweep plot vllm bench sweep plot_pareto vllm bench sweep serve vllm bench sweep serve_sla vllm bench throughput Community Community Contact Us Meetups Sponsors Governance Governance Collaboration Policy Committers Governance Process Blog Forum Slack Table of contents Prerequisites Deploy AnythingLLM¶ AnythingLLM is a full-stack application that enables you to turn any document, resource, or piece of content into context that any LLM can use as references during chatting. It allows you to deploy a large language model (LLM) server with vLLM as the backend, which exposes OpenAI-compatible endpoints. Prerequisites¶ Set up the vLLM environment: pip install vllm Deploy¶ Start the vLLM server with a supported chat-completion model, for example: vllm serve Qwen/Qwen1.5-32B-Chat-AWQ --max-model-len 4096 Download and install AnythingLLM Desktop. Configure the AI provider: At the bottom, click the 🔧 wrench icon -> Open settings -> AI Providers -> LLM. Enter the following values: LLM Provider: Generic OpenAI Base URL: http://{vllm server host}:{vllm server port}/v1 Chat Model Name: Qwen/Qwen1.5-32B-Chat-AWQ Create a workspace: At the bottom, click the ↺ back icon and back to workspaces. Create a workspace (e.g., vllm) and start chatting. Add a document. Click the 📎 attachment icon. Upload a document. Select and move the document into your workspace. Save and embed it. Chat using your document as context. September 11, 2025 ``` pip install vllm ``` **Pattern 4:** Start the vLLM server with a supported chat-completion model, for example: ``` vllm serve Qwen/Qwen1.5-32B-Chat-AWQ --max-model-len 4096 ``` **Pattern 5:** vLLM GitHub Home User Guide User Guide Getting Started Getting Started Quickstart Installation Installation GPU CPU TPU Examples Examples Offline inference Offline inference Async LLM Streaming Audio Language Automatic Prefix Caching Basic Batch LLM Inference Chat With Tools Context Extension Data Parallel Disaggregated Prefill V1 Disaggregated Prefill Encoder Decoder Multimodal KV Load Failure Recovery Test LLM Engine Example LLM Engine Reset Kv Load Sharded State Logits Processor LoRA With Quantization Inference Metrics Mistral-Small MLPSpeculator MultiLoRA Inference Offline Inference with the OpenAI Batch file format Prefix Caching Prompt Embed Inference Qwen2.5-Omni Offline Inference Examples Qwen3 Omni Qwen 1M Reproducibility RLHF RLHF Colocate RLHF Online Quant RLHF Utils Save Sharded State Simple Profiling Skip Loading Weights In Engine Init Spec Decode Structured Outputs Torchrun Dp Example Torchrun Example Vision Language Vision Language Multi Image Online serving Online serving API Client Helm Charts Monitoring Dashboards Disaggregated Encoder Disaggregated Prefill Disaggregated Serving Disaggregated Serving P2P Nccl Xpyd Elastic Ep Gradio OpenAI Chatbot Webserver Gradio Webserver Kv Events Subscriber Multi-Node-Serving Multi Instance Data Parallel OpenAI Chat Completion Client OpenAI Chat Completion Client For Multimodal OpenAI Chat Completion Client With Tools OpenAI Chat Completion Client With Tools Required OpenAI Chat Completion Client With Tools Xlam OpenAI Chat Completion Client With Tools Xlam Streaming OpenAI Chat Completion Tool Calls With Reasoning OpenAI Chat Completion With Reasoning OpenAI Chat Completion With Reasoning Streaming OpenAI Completion Client OpenAI Responses Client OpenAI Responses Client With Mcp Tools OpenAI Responses Client With Tools OpenAI Transcription Client OpenAI Translation Client Setup OpenTelemetry POC Prometheus and Grafana Prompt Embed Inference With OpenAI Client Ray Serve Deepseek Retrieval Augmented Generation With Langchain Retrieval Augmented Generation With Llamaindex Run Cluster Sagemaker-Entrypoint Streamlit OpenAI Chatbot Webserver Structured Outputs Token Generation Client Utils Others Others LMCache Examples Logging Configuration Tensorize vLLM Model Pooling Pooling Classify Embed Plugin Pooling Score Token Classify Token Embed General General vLLM V1 Frequently Asked Questions Production Metrics Reproducibility Security Troubleshooting Usage Stats Collection Inference and Serving Inference and Serving Offline Inference OpenAI-Compatible Server Context Parallel Deployment Data Parallel Deployment Troubleshooting distributed deployments Expert Parallel Deployment Parallelism and Scaling Integrations Integrations LangChain LlamaIndex Deployment Deployment Using Docker Using Kubernetes Using Nginx Frameworks Frameworks Anyscale AnythingLLM AutoGen BentoML Cerebrium Chatbox Dify dstack Haystack Helm Hugging Face Inference Endpoints LiteLLM Lobe Chat LWS Modal Open WebUI Retrieval-Augmented Generation SkyPilot Streamlit NVIDIA Triton Integrations Integrations KAITO KServe Kthena KubeAI KubeRay Llama Stack llm-d llmaz Production stack Training Training Reinforcement Learning from Human Feedback Transformers Reinforcement Learning Configuration Configuration Conserving Memory Engine Arguments Environment Variables Model Resolution Optimization and Tuning Server Arguments TPU Models Models Supported Models Generative Models Pooling Models Extensions Extensions Loading Model weights with fastsafetensors Loading models with Run:ai Model Streamer Loading models with CoreWeave's Tensorizer Hardware Supported Models Hardware Supported Models CPU - Intel® Xeon® XPU - Intel® GPUs TPU Features Features Automatic Prefix Caching Batch Invariance Custom Arguments Custom Logits Processors Disaggregated Encoder Disaggregated Prefilling (experimental) Interleaved Thinking LoRA Adapters MooncakeConnector Usage Guide Multimodal Inputs NixlConnector Usage Guide Prompt Embedding Inputs Reasoning Outputs Sleep Mode Speculative Decoding Structured Outputs Tool Calling Quantization Quantization AutoAWQ AutoRound BitBLAS BitsAndBytes FP8 W8A8 GGUF GPTQModel FP8 INC INT4 W4A16 INT8 W8A8 NVIDIA Model Optimizer Quantized KV Cache AMD Quark TorchAO Developer Guide Developer Guide General General Deprecation Policy Dockerfile Incremental Compilation Workflow Profiling vLLM Vulnerability Management Model Implementation Model Implementation Basic Model Registering a Model Unit Testing Multi-Modal Support Speech-to-Text (Transcription/Translation) Support CI CI CI Failures Nightly Builds of vLLM Wheels Update PyTorch version on vLLM OSS CI/CD Design Documents Design Documents Plugins Plugins IO Processor Plugins LoRA Resolver Plugins Plugin System Architecture Overview Architecture Overview Table of contents Entrypoints LLM Class OpenAI-Compatible API Server LLM Engine LLMEngine AsyncLLMEngine Worker Model Runner Model Class Hierarchy CUDA Graphs Dual Batch Overlap How to debug the vLLM-torch.compile integration Fused MoE Modular Kernel Integration with Hugging Face Hybrid KV Cache Manager Logits Processors Metrics Multi-Modal Data Processing Fused MoE Kernel Features Python Multiprocessing Optimization levels P2P NCCL Connector Paged Attention Automatic Prefix Caching torch.compile integration Benchmarking Benchmarking Benchmark CLI Parameter Sweeps Performance Dashboard API Reference API Reference vllm vllm beam_search collect_env connections env_override envs forward_context logger logits_process logprobs outputs pooling_params sampling_params scalar_type scripts sequence tasks tracing version assets assets audio base image video attention attention layer selector backends backends abstract registry utils layers layers chunked_local_attention cross_attention encoder_only_attention mm_encoder_attention ops ops chunked_prefill_paged_decode common flashmla merge_attn_states paged_attn pallas_kv_cache_update prefix_prefill rocm_aiter_mla_sparse triton_decode_attention triton_merge_attn_states triton_reshape_and_cache_flash triton_unified_attention vit_attn_wrappers utils utils fa_utils kv_sharing_utils kv_transfer_utils benchmarks benchmarks datasets latency serve startup throughput lib lib endpoint_request_func ready_checker utils sweep sweep cli param_sweep plot plot_pareto serve serve_sla server sla_sweep utils compilation compilation activation_quant_fusion backends base_static_graph caching collective_fusion compiler_interface counter cuda_graph decorators fix_functionalization fusion fusion_attn fx_utils inductor_pass matcher_utils monitor noop_elimination partition_rules pass_manager piecewise_backend post_cleanup qk_norm_rope_fusion rocm_aiter_fusion sequence_parallelism torch25_custom_graph_pass vllm_inductor_pass wrapper config config attention cache compilation device ec_transfer kv_events kv_transfer load lora model multimodal observability parallel pooler profiler scheduler speculative speech_to_text structured_outputs utils vllm device_allocator device_allocator cumem distributed distributed communication_op kv_events parallel_state tpu_distributed_utils utils device_communicators device_communicators all2all all_reduce_utils base_device_communicator cpu_communicator cuda_communicator cuda_wrapper custom_all_reduce mnnvl_compat pynccl pynccl_allocator pynccl_wrapper quick_all_reduce ray_communicator shm_broadcast shm_object_storage symm_mem tpu_communicator xpu_communicator ec_transfer ec_transfer ec_transfer_state ec_connector ec_connector base example_connector factory eplb eplb async_worker eplb_state rebalance_execute policy policy abstract default kv_transfer kv_transfer kv_transfer_state kv_connector kv_connector base factory utils v1 v1 base decode_bench_connector example_connector lmcache_connector lmcache_mp_connector metrics mooncake_connector multi_connector nixl_connector offloading_connector lmcache_integration lmcache_integration multi_process_adapter utils vllm_v1_adapter p2p p2p p2p_nccl_connector p2p_nccl_engine tensor_memory_pool engine engine arg_utils async_llm_engine llm_engine protocol entrypoints entrypoints api_server chat_utils constants context launcher llm logger renderer responses_utils score_utils ssl tool tool_server utils anthropic anthropic protocol serving_messages cli cli collect_env main openai run_batch serve types benchmark benchmark base latency main serve startup sweep throughput openai openai api_server cli_args orca_metrics protocol run_batch serving_chat serving_chat_stream_harmony serving_completion serving_engine serving_models serving_responses serving_transcription speech_to_text utils parser parser harmony_utils responses_parser pooling pooling classify classify api_router protocol serving embed embed api_router conftest protocol serving pooling pooling api_router protocol serving score score api_router protocol serving sagemaker sagemaker routes serve serve cache cache api_router disagg disagg api_router protocol serving elastic_ep elastic_ep api_router middleware instrumentator instrumentator health metrics server_info lora lora api_router profile profile api_router rlhf rlhf api_router rpc rpc api_router sleep sleep api_router tokenize tokenize api_router serving inputs inputs data parse preprocess logging_utils logging_utils dump_input formatter lazy log_time lora lora lora_model lora_weights model_manager peft_helper request resolver utils worker_manager layers layers base base_linear column_parallel_linear fused_moe logits_processor replicated_linear row_parallel_linear utils vocal_parallel_embedding ops ops ipex_ops ipex_ops lora_ops torch_ops torch_ops lora_ops triton_ops triton_ops fused_moe_lora_op kernel_utils lora_expand_op lora_kernel_metadata lora_shrink_op utils xla_ops xla_ops lora_ops punica_wrapper punica_wrapper punica_base punica_cpu punica_gpu punica_selector punica_tpu punica_xpu utils model_executor model_executor custom_op parameter utils layers layers activation attention_layer_base batch_invariant conv kda layernorm lightning_attn linear logits_processor mla pooler resampler utils vocab_parallel_embedding fla fla ops ops chunk chunk_delta_h chunk_o chunk_scaled_dot_kkt cumsum fused_recurrent index kda l2norm layernorm_guard op solve_tril utils wy_fast fused_moe fused_moe all2all_utils batched_deep_gemm_moe config cpu_fused_moe cutlass_moe deep_gemm_moe deep_gemm_utils deepep_ht_prepare_finalize deepep_ll_prepare_finalize flashinfer_cutedsl_moe flashinfer_cutlass_moe flashinfer_cutlass_prepare_finalize flashinfer_trtllm_moe fused_batched_moe fused_marlin_moe fused_moe fused_moe_method_base fused_moe_modular_method gpt_oss_triton_kernels_moe layer modular_kernel moe_align_block_size moe_pallas moe_permute_unpermute moe_torch_iterative pplx_prepare_finalize prepare_finalize rocm_aiter_fused_moe routing_simulator shared_fused_moe topk_weight_and_reduce triton_deep_gemm_moe trtllm_moe unquantized_fused_moe_method utils zero_expert_fused_moe mamba mamba abstract linear_attn mamba_mixer mamba_mixer2 mamba_utils short_conv ops ops causal_conv1d layernorm_gated mamba_ssm ssd_bmm ssd_chunk_scan ssd_chunk_state ssd_combined ssd_state_passing quantization quantization auto_round awq awq_marlin awq_triton base_config bitblas bitsandbytes cpu_wna16 deepspeedfp experts_int8 fbgemm_fp8 fp8 fp_quant gguf gptq gptq_bitblas gptq_marlin gptq_marlin_24 hqq_marlin inc input_quant_fp8 ipex_quant kv_cache modelopt moe_wna16 mxfp4 petit ptpc_fp8 qutlass_utils rtn schema torchao tpu_int8 compressed_tensors compressed_tensors compressed_tensors compressed_tensors_moe triton_scaled_mm utils schemes schemes compressed_tensors_24 compressed_tensors_scheme compressed_tensors_w4a4_nvfp4 compressed_tensors_w4a8_fp8 compressed_tensors_w4a8_int compressed_tensors_w4a16_24 compressed_tensors_w4a16_nvfp4 compressed_tensors_w8a8_fp8 compressed_tensors_w8a8_int8 compressed_tensors_w8a16_fp8 compressed_tensors_wNa16 transform transform linear module utils schemes schemes linear_qutlass_nvfp4 kernels kernels mixed_precision mixed_precision allspark bitblas conch cutlass dynamic_4bit exllama MPLinearKernel machete marlin xpu scaled_mm scaled_mm aiter cpu cutlass ScaledMMLinearKernel triton xla quark quark quark quark_moe utils schemes schemes quark_ocp_mx quark_scheme quark_w8a8_fp8 quark_w8a8_int8 utils utils allspark_utils bitblas_utils flashinfer_fp4_moe flashinfer_utils fp8_utils gptq_utils int8_utils layer_utils machete_utils marlin_utils marlin_utils_fp4 marlin_utils_fp8 marlin_utils_test marlin_utils_test_24 mxfp4_utils mxfp6_utils mxfp8_utils nvfp4_emulation_utils nvfp4_moe_support ocp_mx_utils petit_utils quant_utils w8a8_utils rotary_embedding rotary_embedding base common deepseek_scaling_rope dual_chunk_rope dynamic_ntk_alpha_rope dynamic_ntk_scaling_rope ernie45_vl_rope linear_scaling_rope llama3_rope llama4_vision_rope mrope ntk_scaling_rope phi3_long_rope_scaled_rope xdrope yarn_scaling_rope model_loader model_loader base_loader bitsandbytes_loader default_loader dummy_loader gguf_loader online_quantization runai_streamer_loader sharded_state_loader tensorizer tensorizer_loader tpu utils weight_utils models models adapters afmoe aimv2 apertus arcee arctic aria audioflamingo3 aya_vision bagel baichuan bailing_moe bamba bee bert bert_with_rope blip blip2 bloom chameleon chatglm clip cohere2_vision commandr config dbrx deepencoder deepseek_eagle deepseek_mtp deepseek_ocr deepseek_v2 deepseek_vl2 dots1 dots_ocr ernie45 ernie45_moe ernie45_vl ernie45_vl_moe ernie_mtp exaone exaone4 fairseq2_llama falcon falcon_h1 flex_olmo fuyu gemma gemma2 gemma3 gemma3_mm gemma3n gemma3n_mm glm glm4 glm4_1v glm4_moe glm4_moe_mtp glm4v gpt2 gpt_bigcode gpt_j gpt_neox gpt_oss granite granite_speech granitemoe granitemoehybrid granitemoeshared gritlm grok1 h2ovl hunyuan_v1 hunyuan_vision hyperclovax_vision idefics2_vision_model idefics3 interfaces interfaces_base intern_vit internlm2 internlm2_ve interns1 interns1_vit internvl jais jais2 jamba jina_vl keye keye_vl1_5 kimi_linear kimi_vl lfm2 lfm2_moe lightonocr llama llama4 llama4_eagle llama_eagle llama_eagle3 llava llava_next llava_next_video llava_onevision longcat_flash longcat_flash_mtp mamba mamba2 medusa midashenglm mimo mimo_mtp mimo_v2_flash minicpm minicpm3 minicpm_eagle minicpmo minicpmv minimax_m2 minimax_text_01 minimax_vl_01 mistral3 mistral_large_3 mistral_large_3_eagle mixtral mllama4 mlp_speculator modernbert module_mapping molmo moonvit mpt nano_nemotron_vl nemotron nemotron_h nemotron_nas nemotron_vl nvlm_d olmo olmo2 olmoe opencua openpangu openpangu_mtp opt orion ouro ovis ovis2_5 paddleocr_vl paligemma persimmon phi phi3 phi3v phi4mm phi4mm_audio phi4mm_utils phimoe pixtral plamo2 plamo3 qwen qwen2 qwen2_5_omni_thinker qwen2_5_vl qwen2_audio qwen2_moe qwen2_rm qwen2_vl qwen3 qwen3_moe qwen3_next qwen3_next_mtp qwen3_omni_moe_thinker qwen3_vl qwen3_vl_moe qwen_vl radio registry roberta rvl seed_oss siglip siglip2navit skyworkr1v smolvlm solar stablelm starcoder2 step3_text step3_vl swin tarsier telechat2 teleflm terratorch ultravox utils vision voxtral voxtral_streaming whisper whisper_utils zamba2 transformers transformers base causal legacy moe multimodal pooling utils warmup warmup deep_gemm_warmup kernel_warmup multimodal multimodal audio base cache evs hasher image inputs parse processing profiling registry utils video platforms platforms cpu cuda interface rocm tpu xpu plugins plugins io_processors io_processors interface lora_resolvers lora_resolvers filesystem_resolver profiler profiler layerwise_profile utils wrapper ray ray lazy_utils ray_env reasoning reasoning abs_reasoning_parsers basic_parsers deepseek_r1_reasoning_parser deepseek_v3_reasoning_parser ernie45_reasoning_parser glm4_moe_reasoning_parser gptoss_reasoning_parser granite_reasoning_parser holo2_reasoning_parser hunyuan_a13b_reasoning_parser identity_reasoning_parser minimax_m2_reasoning_parser mistral_reasoning_parser olmo3_reasoning_parser qwen3_reasoning_parser seedoss_reasoning_parser step3_reasoning_parser tokenizers tokenizers deepseek_v32 deepseek_v32_encoding detokenizer_utils hf mistral protocol registry tool_parsers tool_parsers abstract_tool_parser deepseekv3_tool_parser deepseekv31_tool_parser deepseekv32_tool_parser ernie45_tool_parser functiongemma_tool_parser gigachat3_tool_parser glm4_moe_tool_parser glm47_moe_tool_parser granite_20b_fc_tool_parser granite_tool_parser hermes_tool_parser hunyuan_a13b_tool_parser internlm2_tool_parser jamba_tool_parser kimi_k2_tool_parser llama4_pythonic_tool_parser llama_tool_parser longcat_tool_parser minimax_m2_tool_parser minimax_tool_parser mistral_tool_parser olmo3_tool_parser openai_tool_parser phi4mini_tool_parser pythonic_tool_parser qwen3coder_tool_parser qwen3xml_tool_parser seed_oss_tool_parser step3_tool_parser utils xlam_tool_parser transformers_utils transformers_utils config config_parser_base dynamic_module gguf_utils processor repo_utils runai_utils s3_utils tokenizer utils chat_templates chat_templates registry configs configs afmoe arctic bagel chatglm deepseek_vl2 dotsocr eagle falcon flex_olmo hunyuan_vl jais kimi_linear kimi_vl lfm2_moe medusa midashenglm mistral mlp_speculator moonvit nemotron nemotron_h olmo3 ovis qwen3_next radio step3_vl tarsier2 ultravox speculators speculators algos base processors processors bagel deepseek_ocr deepseek_vl2 hunyuan_vl hunyuan_vl_image ovis ovis2_5 triton_utils triton_utils importing usage usage usage_lib utils utils argparse_utils async_utils cache collection_utils counter deep_gemm flashinfer func_utils gc_utils hashing import_utils jsontree math_utils mem_constants mem_utils nccl network_utils nvtx_pytorch_hooks platform_utils profiling registry serial_utils system_utils tensor_schema torch_utils v1 v1 cudagraph_dispatcher kv_cache_interface outputs request serial_utils utils attention attention backends backends cpu_attn flash_attn flashinfer flex_attention gdn_attn linear_attn mamba1_attn mamba2_attn mamba_attn pallas rocm_aiter_fa rocm_aiter_unified_attn rocm_attn short_conv_attn tree_attn triton_attn utils mla mla aiter_triton_mla common cutlass_mla flashattn_mla flashinfer_mla flashmla flashmla_sparse indexer rocm_aiter_mla rocm_aiter_mla_sparse triton_mla core core block_pool encoder_cache_manager kv_cache_coordinator kv_cache_manager kv_cache_metrics kv_cache_utils single_type_kv_cache_manager sched sched async_scheduler interface output request_queue scheduler utils engine engine async_llm coordinator core core_client detokenizer exceptions input_processor llm_engine logprobs output_processor parallel_sampling utils executor executor abstract multiproc_executor ray_distributed_executor ray_executor ray_utils uniproc_executor kv_offload kv_offload abstract arc_manager backend cpu factory lru_manager mediums spec backends backends cpu worker worker cpu_gpu worker metrics metrics loggers perf prometheus ray_wrappers reader stats pool pool metadata sample sample metadata rejection_sampler sampler logits_processor logits_processor builtin interface state ops ops bad_words logprobs penalties topk_topp_sampler tpu tpu metadata sampler spec_decode spec_decode eagle medusa metadata metrics ngram_proposer suffix_decoding utils structured_output structured_output backend_guidance backend_lm_format_enforcer backend_outlines backend_types backend_xgrammar request utils worker worker block_table cp_utils cpu_model_runner cpu_worker dp_utils ec_connector_model_runner_mixin gpu_input_batch gpu_model_runner gpu_ubatch_wrapper gpu_worker kv_connector_model_runner_mixin lora_model_runner_mixin tpu_input_batch tpu_model_runner tpu_worker ubatch_utils ubatching utils worker_base workspace xpu_model_runner xpu_worker gpu gpu async_utils attn_utils block_table cudagraph_utils dp_utils input_batch model_runner states structured_outputs metrics metrics logits sample sample gumbel logprob metadata min_p output penalties sampler spec_decode spec_decode eagle eagle_cudagraph rejection_sample CLI Reference CLI Reference vllm serve vllm chat vllm complete vllm run-batch vllm bench vllm bench vllm bench latency vllm bench serve vllm bench sweep plot vllm bench sweep plot_pareto vllm bench sweep serve vllm bench sweep serve_sla vllm bench throughput Community Community Contact Us Meetups Sponsors Governance Governance Collaboration Policy Committers Governance Process Blog Forum Slack Table of contents Entrypoints LLM Class OpenAI-Compatible API Server LLM Engine LLMEngine AsyncLLMEngine Worker Model Runner Model Class Hierarchy Architecture Overview¶ This document provides an overview of the vLLM architecture. Architecture Overview Entrypoints LLM Class OpenAI-Compatible API Server LLM Engine LLMEngine AsyncLLMEngine Worker Model Runner Model Class Hierarchy Entrypoints¶ vLLM provides a number of entrypoints for interacting with the system. The following diagram shows the relationship between them. LLM Class¶ The LLM class provides the primary Python interface for doing offline inference, which is interacting with a model without using a separate model inference server. Here is a sample of LLM class usage: Code from vllm import LLM, SamplingParams # Define a list of input prompts prompts = [ "Hello, my name is", "The capital of France is", "The largest ocean is", ] # Define sampling parameters sampling_params = SamplingParams(temperature=0.8, top_p=0.95) # Initialize the LLM engine with the OPT-125M model llm = LLM(model="facebook/opt-125m") # Generate outputs for the input prompts outputs = llm.generate(prompts, sampling_params) # Print the generated outputs for output in outputs: prompt = output.prompt generated_text = output.outputs[0].text print(f"Prompt: {prompt!r}, Generated text: {generated_text!r}") More API details can be found in the Offline Inference section of the API docs. The code for the LLM class can be found in vllm/entrypoints/llm.py. OpenAI-Compatible API Server¶ The second primary interface to vLLM is via its OpenAI-compatible API server. This server can be started using the vllm serve command. vllm serve The code for the vllm CLI can be found in vllm/entrypoints/cli/main.py. Sometimes you may see the API server entrypoint used directly instead of via the vllm CLI command. For example: python -m vllm.entrypoints.openai.api_server --model Warning python -m vllm.entrypoints.openai.api_server is deprecated and may become unsupported in a future release. That code can be found in vllm/entrypoints/openai/api_server.py. More details on the API server can be found in the OpenAI-Compatible Server document. LLM Engine¶ The LLMEngine and AsyncLLMEngine classes are central to the functioning of the vLLM system, handling model inference and asynchronous request processing. LLMEngine¶ The LLMEngine class is the core component of the vLLM engine. It is responsible for receiving requests from clients and generating outputs from the model. The LLMEngine includes input processing, model execution (possibly distributed across multiple hosts and/or GPUs), scheduling, and output processing. Input Processing: Handles tokenization of input text using the specified tokenizer. Scheduling: Chooses which requests are processed in each step. Model Execution: Manages the execution of the language model, including distributed execution across multiple GPUs. Output Processing: Processes the outputs generated by the model, decoding the token IDs from a language model into human-readable text. The code for LLMEngine can be found in vllm/engine/llm_engine.py. AsyncLLMEngine¶ The AsyncLLMEngine class is an asynchronous wrapper for the LLMEngine class. It uses asyncio to create a background loop that continuously processes incoming requests. The AsyncLLMEngine is designed for online serving, where it can handle multiple concurrent requests and stream outputs to clients. The OpenAI-compatible API server uses the AsyncLLMEngine. There is also a demo API server that serves as a simpler example in vllm/entrypoints/api_server.py. The code for AsyncLLMEngine can be found in vllm/engine/async_llm_engine.py. Worker¶ A worker is a process that runs the model inference. vLLM follows the common practice of using one process to control one accelerator device, such as GPUs. For example, if we use tensor parallelism of size 2 and pipeline parallelism of size 2, we will have 4 workers in total. Workers are identified by their rank and local_rank. rank is used for global orchestration, while local_rank is mainly used for assigning the accelerator device and accessing local resources such as the file system and shared memory. Model Runner¶ Every worker has one model runner object, responsible for loading and running the model. Much of the model execution logic resides here, such as preparing input tensors and capturing cudagraphs. Model¶ Every model runner object has one model object, which is the actual torch.nn.Module instance. See huggingface_integration for how various configurations affect the class we ultimately get. Class Hierarchy¶ The following figure shows the class hierarchy of vLLM: There are several important design choices behind this class hierarchy: 1. Extensibility: All classes in the hierarchy accept a configuration object containing all the necessary information. The VllmConfig class is the main configuration object that is passed around. The class hierarchy is quite deep, and every class needs to read the configuration it is interested in. By encapsulating all configurations in one object, we can easily pass the configuration object around and access the configuration we need. Suppose we want to add a new feature (this is often the case given how fast the field of LLM inference is evolving) that only touches the model runner. We will have to add a new configuration option in the VllmConfig class. Since we pass the whole config object around, we only need to add the configuration option to the VllmConfig class, and the model runner can access it directly. We don't need to change the constructor of the engine, worker, or model class to pass the new configuration option. 2. Uniformity: The model runner needs a unified interface to create and initialize the model. vLLM supports more than 50 types of popular open-source models. Each model has its own initialization logic. If the constructor signature varies with models, the model runner does not know how to call the constructor accordingly, without complicated and error-prone inspection logic. By making the constructor of the model class uniform, the model runner can easily create and initialize the model without knowing the specific model type. This is also useful for composing models. Vision-language models often consist of a vision model and a language model. By making the constructor uniform, we can easily create a vision model and a language model and compose them into a vision-language model. Note To support this change, all vLLM models' signatures have been updated to: def __init__(self, *, vllm_config: VllmConfig, prefix: str = ""): To avoid accidentally passing incorrect arguments, the constructor is now keyword-only. This ensures that the constructor will raise an error if old configurations are passed. vLLM developers have already made this change for all models within vLLM. For out-of-tree registered models, developers need to update their models, for example by adding shim code to adapt the old constructor signature to the new one: Code class MyOldModel(nn.Module): def __init__( self, config, cache_config: Optional[CacheConfig] = None, quant_config: Optional[QuantizationConfig] = None, lora_config: Optional[LoRAConfig] = None, prefix: str = "", ) -> None: ... from vllm.config import VllmConfig class MyNewModel(MyOldModel): def __init__(self, *, vllm_config: VllmConfig, prefix: str = ""): config = vllm_config.model_config.hf_config cache_config = vllm_config.cache_config quant_config = vllm_config.quant_config lora_config = vllm_config.lora_config super().__init__(config, cache_config, quant_config, lora_config, prefix) from packaging import version if version.parse(__version__) >= version.parse("0.6.4"): MyModel = MyNewModel else: MyModel = MyOldModel This way, the model can work with both old and new versions of vLLM. 3. Sharding and Quantization at Initialization: Certain features require changing the model weights. For example, tensor parallelism needs to shard the model weights, and quantization needs to quantize the model weights. There are two possible ways to implement this feature. One way is to change the model weights after the model is initialized. The other way is to change the model weights during the model initialization. vLLM chooses the latter. The first approach is not scalable to large models. Suppose we want to run a 405B model (with roughly 810GB weights) with 16 H100 80GB GPUs. Ideally, every GPU should only load 50GB weights. If we change the model weights after the model is initialized, we need to load the full 810GB weights to every GPU and then shard the weights, leading to a huge memory overhead. Instead, if we shard the weights during the model initialization, every layer will only create a shard of the weights it needs, leading to a much smaller memory overhead. The same idea applies to quantization. Note that we also add an additional argument prefix to the model's constructor so that the model can initialize itself differently based on the prefix. This is useful for non-uniform quantization, where different parts of the model are quantized differently. The prefix is usually an empty string for the top-level model and a string like "vision" or "language" for the sub-models. In general, it matches the name of the module's state dict in the checkpoint file. One disadvantage of this design is that it is hard to write unit tests for individual components in vLLM because every component needs to be initialized by a complete config object. We solve this problem by providing a default initialization function that creates a default config object with all fields set to None. If the component we want to test only cares about a few fields in the config object, we can create a default config object and set the fields we care about. This way, we can test the component in isolation. Note that many tests in vLLM are end-to-end tests that test the whole system, so this is not a big problem. In summary, the complete config object VllmConfig can be treated as an engine-level global state that is shared among all vLLM classes. October 17, 2025 ``` LLM ``` **Pattern 6:** Sometimes you may see the API server entrypoint used directly instead of via the vllm CLI command. For example: ``` vllm ``` **Pattern 7:** vLLM GitHub Home User Guide User Guide Getting Started Getting Started Quickstart Installation Installation GPU CPU TPU Examples Examples Offline inference Offline inference Async LLM Streaming Audio Language Automatic Prefix Caching Basic Batch LLM Inference Chat With Tools Context Extension Data Parallel Disaggregated Prefill V1 Disaggregated Prefill Encoder Decoder Multimodal KV Load Failure Recovery Test LLM Engine Example LLM Engine Reset Kv Load Sharded State Logits Processor LoRA With Quantization Inference Metrics Mistral-Small MLPSpeculator MultiLoRA Inference Offline Inference with the OpenAI Batch file format Prefix Caching Prompt Embed Inference Qwen2.5-Omni Offline Inference Examples Qwen3 Omni Qwen 1M Reproducibility RLHF RLHF Colocate RLHF Online Quant RLHF Utils Save Sharded State Simple Profiling Skip Loading Weights In Engine Init Spec Decode Structured Outputs Torchrun Dp Example Torchrun Example Vision Language Vision Language Multi Image Online serving Online serving API Client Helm Charts Monitoring Dashboards Disaggregated Encoder Disaggregated Prefill Disaggregated Serving Disaggregated Serving P2P Nccl Xpyd Elastic Ep Gradio OpenAI Chatbot Webserver Gradio Webserver Kv Events Subscriber Multi-Node-Serving Multi Instance Data Parallel OpenAI Chat Completion Client OpenAI Chat Completion Client For Multimodal OpenAI Chat Completion Client With Tools OpenAI Chat Completion Client With Tools Required OpenAI Chat Completion Client With Tools Xlam OpenAI Chat Completion Client With Tools Xlam Streaming OpenAI Chat Completion Tool Calls With Reasoning OpenAI Chat Completion With Reasoning OpenAI Chat Completion With Reasoning Streaming OpenAI Completion Client OpenAI Responses Client OpenAI Responses Client With Mcp Tools OpenAI Responses Client With Tools OpenAI Transcription Client OpenAI Translation Client Setup OpenTelemetry POC Prometheus and Grafana Prompt Embed Inference With OpenAI Client Ray Serve Deepseek Retrieval Augmented Generation With Langchain Retrieval Augmented Generation With Llamaindex Run Cluster Sagemaker-Entrypoint Streamlit OpenAI Chatbot Webserver Structured Outputs Token Generation Client Utils Others Others LMCache Examples Logging Configuration Tensorize vLLM Model Pooling Pooling Classify Embed Plugin Pooling Score Token Classify Token Embed General General vLLM V1 Frequently Asked Questions Production Metrics Reproducibility Security Troubleshooting Usage Stats Collection Inference and Serving Inference and Serving Offline Inference OpenAI-Compatible Server Context Parallel Deployment Data Parallel Deployment Troubleshooting distributed deployments Expert Parallel Deployment Parallelism and Scaling Integrations Integrations LangChain LlamaIndex Deployment Deployment Using Docker Using Kubernetes Using Nginx Frameworks Frameworks Anyscale AnythingLLM AutoGen BentoML Cerebrium Chatbox Dify dstack Haystack Helm Hugging Face Inference Endpoints LiteLLM Lobe Chat LWS Modal Open WebUI Retrieval-Augmented Generation SkyPilot Streamlit NVIDIA Triton Integrations Integrations KAITO KServe Kthena KubeAI KubeRay Llama Stack llm-d llmaz Production stack Training Training Reinforcement Learning from Human Feedback Transformers Reinforcement Learning Configuration Configuration Conserving Memory Engine Arguments Environment Variables Model Resolution Optimization and Tuning Server Arguments TPU Models Models Supported Models Generative Models Pooling Models Extensions Extensions Loading Model weights with fastsafetensors Loading models with Run:ai Model Streamer Loading models with CoreWeave's Tensorizer Hardware Supported Models Hardware Supported Models CPU - Intel® Xeon® XPU - Intel® GPUs TPU Features Features Automatic Prefix Caching Batch Invariance Custom Arguments Custom Logits Processors Disaggregated Encoder Disaggregated Prefilling (experimental) Interleaved Thinking LoRA Adapters MooncakeConnector Usage Guide Multimodal Inputs NixlConnector Usage Guide Prompt Embedding Inputs Reasoning Outputs Sleep Mode Speculative Decoding Structured Outputs Tool Calling Quantization Quantization AutoAWQ AutoRound BitBLAS BitsAndBytes FP8 W8A8 GGUF GPTQModel FP8 INC INT4 W4A16 INT8 W8A8 NVIDIA Model Optimizer Quantized KV Cache AMD Quark TorchAO Developer Guide Developer Guide General General Deprecation Policy Dockerfile Incremental Compilation Workflow Profiling vLLM Vulnerability Management Model Implementation Model Implementation Basic Model Registering a Model Unit Testing Multi-Modal Support Speech-to-Text (Transcription/Translation) Support CI CI CI Failures Nightly Builds of vLLM Wheels Update PyTorch version on vLLM OSS CI/CD Design Documents Design Documents Plugins Plugins IO Processor Plugins LoRA Resolver Plugins Plugin System Architecture Overview CUDA Graphs Dual Batch Overlap How to debug the vLLM-torch.compile integration Fused MoE Modular Kernel Integration with Hugging Face Hybrid KV Cache Manager Logits Processors Metrics Multi-Modal Data Processing Fused MoE Kernel Features Python Multiprocessing Optimization levels P2P NCCL Connector Paged Attention Automatic Prefix Caching torch.compile integration Benchmarking Benchmarking Benchmark CLI Parameter Sweeps Performance Dashboard API Reference API Reference vllm vllm beam_search collect_env connections env_override envs forward_context logger logits_process logprobs outputs pooling_params sampling_params scalar_type scripts sequence tasks tracing version assets assets audio base image video attention attention layer selector backends backends abstract registry utils layers layers chunked_local_attention cross_attention encoder_only_attention mm_encoder_attention ops ops chunked_prefill_paged_decode common flashmla merge_attn_states paged_attn pallas_kv_cache_update prefix_prefill rocm_aiter_mla_sparse triton_decode_attention triton_merge_attn_states triton_reshape_and_cache_flash triton_unified_attention vit_attn_wrappers utils utils fa_utils kv_sharing_utils kv_transfer_utils benchmarks benchmarks datasets latency serve startup throughput lib lib endpoint_request_func ready_checker utils sweep sweep cli param_sweep plot plot_pareto serve serve_sla server sla_sweep utils compilation compilation activation_quant_fusion backends base_static_graph caching collective_fusion compiler_interface counter cuda_graph decorators fix_functionalization fusion fusion_attn fx_utils inductor_pass matcher_utils monitor noop_elimination partition_rules pass_manager piecewise_backend post_cleanup qk_norm_rope_fusion rocm_aiter_fusion sequence_parallelism torch25_custom_graph_pass vllm_inductor_pass wrapper config config attention cache compilation device ec_transfer kv_events kv_transfer load lora model multimodal observability parallel pooler profiler scheduler speculative speech_to_text structured_outputs utils vllm device_allocator device_allocator cumem distributed distributed communication_op kv_events parallel_state tpu_distributed_utils utils device_communicators device_communicators all2all all_reduce_utils base_device_communicator cpu_communicator cuda_communicator cuda_wrapper custom_all_reduce mnnvl_compat pynccl pynccl_allocator pynccl_wrapper quick_all_reduce ray_communicator shm_broadcast shm_object_storage symm_mem tpu_communicator xpu_communicator ec_transfer ec_transfer ec_transfer_state ec_connector ec_connector base example_connector factory eplb eplb async_worker eplb_state rebalance_execute policy policy abstract default kv_transfer kv_transfer kv_transfer_state kv_connector kv_connector base factory utils v1 v1 base decode_bench_connector example_connector lmcache_connector lmcache_mp_connector metrics mooncake_connector multi_connector nixl_connector offloading_connector lmcache_integration lmcache_integration multi_process_adapter utils vllm_v1_adapter p2p p2p p2p_nccl_connector p2p_nccl_engine tensor_memory_pool engine engine arg_utils arg_utils Table of contents NEEDS_HELP T TypeHint TypeHintT logger AsyncEngineArgs enable_log_requests __init__ add_cli_args EngineArgs _api_process_count _api_process_rank additional_config aggregate_engine_logging all2all_backend allowed_local_media_path allowed_media_domains async_scheduling attention_backend attention_config block_size calculate_kv_scales code_revision collect_detailed_traces compilation_config config_format convert cp_kv_cache_interleave_size cpu_offload_gb cudagraph_capture_sizes cudagraph_metrics data_parallel_address data_parallel_backend data_parallel_external_lb data_parallel_hybrid_lb data_parallel_rank data_parallel_rpc_port data_parallel_size data_parallel_size_local data_parallel_start_rank dbo_decode_token_threshold dbo_prefill_token_threshold dcp_kv_cache_interleave_size decode_context_parallel_size default_mm_loras disable_cascade_attn disable_chunked_mm_input disable_custom_all_reduce disable_hybrid_kv_cache_manager disable_log_stats disable_nccl_for_dp_synchronization disable_sliding_window distributed_executor_backend download_dir dtype ec_transfer_config enable_chunked_prefill enable_dbo enable_eplb enable_expert_parallel enable_layerwise_nvtx_tracing enable_lora enable_mfu_metrics enable_mm_embeds enable_prefix_caching enable_prompt_embeds enable_sleep_mode enforce_eager eplb_config expert_placement_strategy fully_sharded_loras generation_config gpu_memory_utilization hf_config_path hf_overrides hf_token ignore_patterns interleave_mm_strings io_processor_plugin kv_cache_dtype kv_cache_memory_bytes kv_cache_metrics kv_cache_metrics_sample kv_events_config kv_offloading_backend kv_offloading_size kv_sharing_fast_prefill kv_transfer_config limit_mm_per_prompt load_format logits_processor_pattern logits_processors logprobs_mode long_prefill_token_threshold lora_dtype mamba_block_size mamba_cache_dtype mamba_ssm_cache_dtype master_addr master_port max_cpu_loras max_cudagraph_capture_size max_logprobs max_long_partial_prefills max_lora_rank max_loras max_model_len max_num_batched_tokens max_num_partial_prefills max_num_seqs max_parallel_loading_workers media_io_kwargs mm_encoder_attn_backend mm_encoder_tp_mode mm_processor_cache_gb mm_processor_cache_type mm_processor_kwargs mm_shm_cache_max_object_size_mb model model_impl model_loader_extra_config nnodes node_rank num_gpu_blocks_override optimization_level otlp_traces_endpoint override_attention_dtype override_generation_config pipeline_parallel_size pooler_config prefill_context_parallel_size prefix_caching_hash_algo profiler_config pt_load_map_location quantization ray_workers_use_nsight reasoning_parser reasoning_parser_plugin revision runner safetensors_load_strategy scheduler_cls scheduling_policy seed served_model_name show_hidden_metrics_for_version skip_mm_profiling skip_tokenizer_init speculative_config stream_interval structured_outputs_config swap_space tensor_parallel_size tokenizer tokenizer_mode tokenizer_revision tokens_only trust_remote_code ubatch_size use_tqdm_on_load video_pruning_rate worker_cls worker_extension_cls __init__ __post_init__ _check_feature_supported _set_default_chunked_prefill_and_prefix_caching_args _set_default_max_num_seqs_and_batched_tokens_args add_cli_args create_engine_config create_load_config create_model_config create_speculative_config from_cli_args get_batch_defaults validate_tensorizer_args _compute_kwargs _raise_unsupported_error collection_to_kwargs contains_type get_kwargs get_type get_type_hints human_readable_int human_readable_int_or_auto is_not_builtin is_online_quantization is_type literal_to_kwargs optional_type parse_type union_dict_and_str async_llm_engine llm_engine protocol entrypoints entrypoints api_server chat_utils constants context launcher llm logger renderer responses_utils score_utils ssl tool tool_server utils anthropic anthropic protocol serving_messages cli cli collect_env main openai run_batch serve types benchmark benchmark base latency main serve startup sweep throughput openai openai api_server cli_args orca_metrics protocol run_batch serving_chat serving_chat_stream_harmony serving_completion serving_engine serving_models serving_responses serving_transcription speech_to_text utils parser parser harmony_utils responses_parser pooling pooling classify classify api_router protocol serving embed embed api_router conftest protocol serving pooling pooling api_router protocol serving score score api_router protocol serving sagemaker sagemaker routes serve serve cache cache api_router disagg disagg api_router protocol serving elastic_ep elastic_ep api_router middleware instrumentator instrumentator health metrics server_info lora lora api_router profile profile api_router rlhf rlhf api_router rpc rpc api_router sleep sleep api_router tokenize tokenize api_router serving inputs inputs data parse preprocess logging_utils logging_utils dump_input formatter lazy log_time lora lora lora_model lora_weights model_manager peft_helper request resolver utils worker_manager layers layers base base_linear column_parallel_linear fused_moe logits_processor replicated_linear row_parallel_linear utils vocal_parallel_embedding ops ops ipex_ops ipex_ops lora_ops torch_ops torch_ops lora_ops triton_ops triton_ops fused_moe_lora_op kernel_utils lora_expand_op lora_kernel_metadata lora_shrink_op utils xla_ops xla_ops lora_ops punica_wrapper punica_wrapper punica_base punica_cpu punica_gpu punica_selector punica_tpu punica_xpu utils model_executor model_executor custom_op parameter utils layers layers activation attention_layer_base batch_invariant conv kda layernorm lightning_attn linear logits_processor mla pooler resampler utils vocab_parallel_embedding fla fla ops ops chunk chunk_delta_h chunk_o chunk_scaled_dot_kkt cumsum fused_recurrent index kda l2norm layernorm_guard op solve_tril utils wy_fast fused_moe fused_moe all2all_utils batched_deep_gemm_moe config cpu_fused_moe cutlass_moe deep_gemm_moe deep_gemm_utils deepep_ht_prepare_finalize deepep_ll_prepare_finalize flashinfer_cutedsl_moe flashinfer_cutlass_moe flashinfer_cutlass_prepare_finalize flashinfer_trtllm_moe fused_batched_moe fused_marlin_moe fused_moe fused_moe_method_base fused_moe_modular_method gpt_oss_triton_kernels_moe layer modular_kernel moe_align_block_size moe_pallas moe_permute_unpermute moe_torch_iterative pplx_prepare_finalize prepare_finalize rocm_aiter_fused_moe routing_simulator shared_fused_moe topk_weight_and_reduce triton_deep_gemm_moe trtllm_moe unquantized_fused_moe_method utils zero_expert_fused_moe mamba mamba abstract linear_attn mamba_mixer mamba_mixer2 mamba_utils short_conv ops ops causal_conv1d layernorm_gated mamba_ssm ssd_bmm ssd_chunk_scan ssd_chunk_state ssd_combined ssd_state_passing quantization quantization auto_round awq awq_marlin awq_triton base_config bitblas bitsandbytes cpu_wna16 deepspeedfp experts_int8 fbgemm_fp8 fp8 fp_quant gguf gptq gptq_bitblas gptq_marlin gptq_marlin_24 hqq_marlin inc input_quant_fp8 ipex_quant kv_cache modelopt moe_wna16 mxfp4 petit ptpc_fp8 qutlass_utils rtn schema torchao tpu_int8 compressed_tensors compressed_tensors compressed_tensors compressed_tensors_moe triton_scaled_mm utils schemes schemes compressed_tensors_24 compressed_tensors_scheme compressed_tensors_w4a4_nvfp4 compressed_tensors_w4a8_fp8 compressed_tensors_w4a8_int compressed_tensors_w4a16_24 compressed_tensors_w4a16_nvfp4 compressed_tensors_w8a8_fp8 compressed_tensors_w8a8_int8 compressed_tensors_w8a16_fp8 compressed_tensors_wNa16 transform transform linear module utils schemes schemes linear_qutlass_nvfp4 kernels kernels mixed_precision mixed_precision allspark bitblas conch cutlass dynamic_4bit exllama MPLinearKernel machete marlin xpu scaled_mm scaled_mm aiter cpu cutlass ScaledMMLinearKernel triton xla quark quark quark quark_moe utils schemes schemes quark_ocp_mx quark_scheme quark_w8a8_fp8 quark_w8a8_int8 utils utils allspark_utils bitblas_utils flashinfer_fp4_moe flashinfer_utils fp8_utils gptq_utils int8_utils layer_utils machete_utils marlin_utils marlin_utils_fp4 marlin_utils_fp8 marlin_utils_test marlin_utils_test_24 mxfp4_utils mxfp6_utils mxfp8_utils nvfp4_emulation_utils nvfp4_moe_support ocp_mx_utils petit_utils quant_utils w8a8_utils rotary_embedding rotary_embedding base common deepseek_scaling_rope dual_chunk_rope dynamic_ntk_alpha_rope dynamic_ntk_scaling_rope ernie45_vl_rope linear_scaling_rope llama3_rope llama4_vision_rope mrope ntk_scaling_rope phi3_long_rope_scaled_rope xdrope yarn_scaling_rope model_loader model_loader base_loader bitsandbytes_loader default_loader dummy_loader gguf_loader online_quantization runai_streamer_loader sharded_state_loader tensorizer tensorizer_loader tpu utils weight_utils models models adapters afmoe aimv2 apertus arcee arctic aria audioflamingo3 aya_vision bagel baichuan bailing_moe bamba bee bert bert_with_rope blip blip2 bloom chameleon chatglm clip cohere2_vision commandr config dbrx deepencoder deepseek_eagle deepseek_mtp deepseek_ocr deepseek_v2 deepseek_vl2 dots1 dots_ocr ernie45 ernie45_moe ernie45_vl ernie45_vl_moe ernie_mtp exaone exaone4 fairseq2_llama falcon falcon_h1 flex_olmo fuyu gemma gemma2 gemma3 gemma3_mm gemma3n gemma3n_mm glm glm4 glm4_1v glm4_moe glm4_moe_mtp glm4v gpt2 gpt_bigcode gpt_j gpt_neox gpt_oss granite granite_speech granitemoe granitemoehybrid granitemoeshared gritlm grok1 h2ovl hunyuan_v1 hunyuan_vision hyperclovax_vision idefics2_vision_model idefics3 interfaces interfaces_base intern_vit internlm2 internlm2_ve interns1 interns1_vit internvl jais jais2 jamba jina_vl keye keye_vl1_5 kimi_linear kimi_vl lfm2 lfm2_moe lightonocr llama llama4 llama4_eagle llama_eagle llama_eagle3 llava llava_next llava_next_video llava_onevision longcat_flash longcat_flash_mtp mamba mamba2 medusa midashenglm mimo mimo_mtp mimo_v2_flash minicpm minicpm3 minicpm_eagle minicpmo minicpmv minimax_m2 minimax_text_01 minimax_vl_01 mistral3 mistral_large_3 mistral_large_3_eagle mixtral mllama4 mlp_speculator modernbert module_mapping molmo moonvit mpt nano_nemotron_vl nemotron nemotron_h nemotron_nas nemotron_vl nvlm_d olmo olmo2 olmoe opencua openpangu openpangu_mtp opt orion ouro ovis ovis2_5 paddleocr_vl paligemma persimmon phi phi3 phi3v phi4mm phi4mm_audio phi4mm_utils phimoe pixtral plamo2 plamo3 qwen qwen2 qwen2_5_omni_thinker qwen2_5_vl qwen2_audio qwen2_moe qwen2_rm qwen2_vl qwen3 qwen3_moe qwen3_next qwen3_next_mtp qwen3_omni_moe_thinker qwen3_vl qwen3_vl_moe qwen_vl radio registry roberta rvl seed_oss siglip siglip2navit skyworkr1v smolvlm solar stablelm starcoder2 step3_text step3_vl swin tarsier telechat2 teleflm terratorch ultravox utils vision voxtral voxtral_streaming whisper whisper_utils zamba2 transformers transformers base causal legacy moe multimodal pooling utils warmup warmup deep_gemm_warmup kernel_warmup multimodal multimodal audio base cache evs hasher image inputs parse processing profiling registry utils video platforms platforms cpu cuda interface rocm tpu xpu plugins plugins io_processors io_processors interface lora_resolvers lora_resolvers filesystem_resolver profiler profiler layerwise_profile utils wrapper ray ray lazy_utils ray_env reasoning reasoning abs_reasoning_parsers basic_parsers deepseek_r1_reasoning_parser deepseek_v3_reasoning_parser ernie45_reasoning_parser glm4_moe_reasoning_parser gptoss_reasoning_parser granite_reasoning_parser holo2_reasoning_parser hunyuan_a13b_reasoning_parser identity_reasoning_parser minimax_m2_reasoning_parser mistral_reasoning_parser olmo3_reasoning_parser qwen3_reasoning_parser seedoss_reasoning_parser step3_reasoning_parser tokenizers tokenizers deepseek_v32 deepseek_v32_encoding detokenizer_utils hf mistral protocol registry tool_parsers tool_parsers abstract_tool_parser deepseekv3_tool_parser deepseekv31_tool_parser deepseekv32_tool_parser ernie45_tool_parser functiongemma_tool_parser gigachat3_tool_parser glm4_moe_tool_parser glm47_moe_tool_parser granite_20b_fc_tool_parser granite_tool_parser hermes_tool_parser hunyuan_a13b_tool_parser internlm2_tool_parser jamba_tool_parser kimi_k2_tool_parser llama4_pythonic_tool_parser llama_tool_parser longcat_tool_parser minimax_m2_tool_parser minimax_tool_parser mistral_tool_parser olmo3_tool_parser openai_tool_parser phi4mini_tool_parser pythonic_tool_parser qwen3coder_tool_parser qwen3xml_tool_parser seed_oss_tool_parser step3_tool_parser utils xlam_tool_parser transformers_utils transformers_utils config config_parser_base dynamic_module gguf_utils processor repo_utils runai_utils s3_utils tokenizer utils chat_templates chat_templates registry configs configs afmoe arctic bagel chatglm deepseek_vl2 dotsocr eagle falcon flex_olmo hunyuan_vl jais kimi_linear kimi_vl lfm2_moe medusa midashenglm mistral mlp_speculator moonvit nemotron nemotron_h olmo3 ovis qwen3_next radio step3_vl tarsier2 ultravox speculators speculators algos base processors processors bagel deepseek_ocr deepseek_vl2 hunyuan_vl hunyuan_vl_image ovis ovis2_5 triton_utils triton_utils importing usage usage usage_lib utils utils argparse_utils async_utils cache collection_utils counter deep_gemm flashinfer func_utils gc_utils hashing import_utils jsontree math_utils mem_constants mem_utils nccl network_utils nvtx_pytorch_hooks platform_utils profiling registry serial_utils system_utils tensor_schema torch_utils v1 v1 cudagraph_dispatcher kv_cache_interface outputs request serial_utils utils attention attention backends backends cpu_attn flash_attn flashinfer flex_attention gdn_attn linear_attn mamba1_attn mamba2_attn mamba_attn pallas rocm_aiter_fa rocm_aiter_unified_attn rocm_attn short_conv_attn tree_attn triton_attn utils mla mla aiter_triton_mla common cutlass_mla flashattn_mla flashinfer_mla flashmla flashmla_sparse indexer rocm_aiter_mla rocm_aiter_mla_sparse triton_mla core core block_pool encoder_cache_manager kv_cache_coordinator kv_cache_manager kv_cache_metrics kv_cache_utils single_type_kv_cache_manager sched sched async_scheduler interface output request_queue scheduler utils engine engine async_llm coordinator core core_client detokenizer exceptions input_processor llm_engine logprobs output_processor parallel_sampling utils executor executor abstract multiproc_executor ray_distributed_executor ray_executor ray_utils uniproc_executor kv_offload kv_offload abstract arc_manager backend cpu factory lru_manager mediums spec backends backends cpu worker worker cpu_gpu worker metrics metrics loggers perf prometheus ray_wrappers reader stats pool pool metadata sample sample metadata rejection_sampler sampler logits_processor logits_processor builtin interface state ops ops bad_words logprobs penalties topk_topp_sampler tpu tpu metadata sampler spec_decode spec_decode eagle medusa metadata metrics ngram_proposer suffix_decoding utils structured_output structured_output backend_guidance backend_lm_format_enforcer backend_outlines backend_types backend_xgrammar request utils worker worker block_table cp_utils cpu_model_runner cpu_worker dp_utils ec_connector_model_runner_mixin gpu_input_batch gpu_model_runner gpu_ubatch_wrapper gpu_worker kv_connector_model_runner_mixin lora_model_runner_mixin tpu_input_batch tpu_model_runner tpu_worker ubatch_utils ubatching utils worker_base workspace xpu_model_runner xpu_worker gpu gpu async_utils attn_utils block_table cudagraph_utils dp_utils input_batch model_runner states structured_outputs metrics metrics logits sample sample gumbel logprob metadata min_p output penalties sampler spec_decode spec_decode eagle eagle_cudagraph rejection_sample CLI Reference CLI Reference vllm serve vllm chat vllm complete vllm run-batch vllm bench vllm bench vllm bench latency vllm bench serve vllm bench sweep plot vllm bench sweep plot_pareto vllm bench sweep serve vllm bench sweep serve_sla vllm bench throughput Community Community Contact Us Meetups Sponsors Governance Governance Collaboration Policy Committers Governance Process Blog Forum Slack Table of contents NEEDS_HELP T TypeHint TypeHintT logger AsyncEngineArgs enable_log_requests __init__ add_cli_args EngineArgs _api_process_count _api_process_rank additional_config aggregate_engine_logging all2all_backend allowed_local_media_path allowed_media_domains async_scheduling attention_backend attention_config block_size calculate_kv_scales code_revision collect_detailed_traces compilation_config config_format convert cp_kv_cache_interleave_size cpu_offload_gb cudagraph_capture_sizes cudagraph_metrics data_parallel_address data_parallel_backend data_parallel_external_lb data_parallel_hybrid_lb data_parallel_rank data_parallel_rpc_port data_parallel_size data_parallel_size_local data_parallel_start_rank dbo_decode_token_threshold dbo_prefill_token_threshold dcp_kv_cache_interleave_size decode_context_parallel_size default_mm_loras disable_cascade_attn disable_chunked_mm_input disable_custom_all_reduce disable_hybrid_kv_cache_manager disable_log_stats disable_nccl_for_dp_synchronization disable_sliding_window distributed_executor_backend download_dir dtype ec_transfer_config enable_chunked_prefill enable_dbo enable_eplb enable_expert_parallel enable_layerwise_nvtx_tracing enable_lora enable_mfu_metrics enable_mm_embeds enable_prefix_caching enable_prompt_embeds enable_sleep_mode enforce_eager eplb_config expert_placement_strategy fully_sharded_loras generation_config gpu_memory_utilization hf_config_path hf_overrides hf_token ignore_patterns interleave_mm_strings io_processor_plugin kv_cache_dtype kv_cache_memory_bytes kv_cache_metrics kv_cache_metrics_sample kv_events_config kv_offloading_backend kv_offloading_size kv_sharing_fast_prefill kv_transfer_config limit_mm_per_prompt load_format logits_processor_pattern logits_processors logprobs_mode long_prefill_token_threshold lora_dtype mamba_block_size mamba_cache_dtype mamba_ssm_cache_dtype master_addr master_port max_cpu_loras max_cudagraph_capture_size max_logprobs max_long_partial_prefills max_lora_rank max_loras max_model_len max_num_batched_tokens max_num_partial_prefills max_num_seqs max_parallel_loading_workers media_io_kwargs mm_encoder_attn_backend mm_encoder_tp_mode mm_processor_cache_gb mm_processor_cache_type mm_processor_kwargs mm_shm_cache_max_object_size_mb model model_impl model_loader_extra_config nnodes node_rank num_gpu_blocks_override optimization_level otlp_traces_endpoint override_attention_dtype override_generation_config pipeline_parallel_size pooler_config prefill_context_parallel_size prefix_caching_hash_algo profiler_config pt_load_map_location quantization ray_workers_use_nsight reasoning_parser reasoning_parser_plugin revision runner safetensors_load_strategy scheduler_cls scheduling_policy seed served_model_name show_hidden_metrics_for_version skip_mm_profiling skip_tokenizer_init speculative_config stream_interval structured_outputs_config swap_space tensor_parallel_size tokenizer tokenizer_mode tokenizer_revision tokens_only trust_remote_code ubatch_size use_tqdm_on_load video_pruning_rate worker_cls worker_extension_cls __init__ __post_init__ _check_feature_supported _set_default_chunked_prefill_and_prefix_caching_args _set_default_max_num_seqs_and_batched_tokens_args add_cli_args create_engine_config create_load_config create_model_config create_speculative_config from_cli_args get_batch_defaults validate_tensorizer_args _compute_kwargs _raise_unsupported_error collection_to_kwargs contains_type get_kwargs get_type get_type_hints human_readable_int human_readable_int_or_auto is_not_builtin is_online_quantization is_type literal_to_kwargs optional_type parse_type union_dict_and_str vllm.engine.arg_utils ¶ NEEDS_HELP module-attribute ¶ NEEDS_HELP = ( any(("--help" in arg) for arg in (argv)) or endswith("mkdocs") or endswith("mkdocs/__main__.py") ) T module-attribute ¶ T = TypeVar('T') TypeHint module-attribute ¶ TypeHint: TypeAlias = type[Any] | object TypeHintT module-attribute ¶ TypeHintT: TypeAlias = type[T] | object logger module-attribute ¶ logger = init_logger(__name__) AsyncEngineArgs dataclass ¶ Bases: EngineArgs Arguments for asynchronous vLLM engine. Source code in vllm/engine/arg_utils.py 2002 2003 2004 2005 2006 2007 2008 2009 2010 2011 2012 2013 2014 2015 2016 2017 2018 2019 2020 2021 2022 2023 2024 2025 2026 2027 2028 2029 2030 2031 2032@dataclass class AsyncEngineArgs(EngineArgs): """Arguments for asynchronous vLLM engine.""" enable_log_requests: bool = False @staticmethod def add_cli_args( parser: FlexibleArgumentParser, async_args_only: bool = False ) -> FlexibleArgumentParser: # Initialize plugin to update the parser, for example, The plugin may # add a new kind of quantization method to --quantization argument or # a new device to --device argument. load_general_plugins() if not async_args_only: parser = EngineArgs.add_cli_args(parser) parser.add_argument( "--enable-log-requests", action=argparse.BooleanOptionalAction, default=AsyncEngineArgs.enable_log_requests, help="Enable logging requests.", ) parser.add_argument( "--disable-log-requests", action=argparse.BooleanOptionalAction, default=not AsyncEngineArgs.enable_log_requests, help="[DEPRECATED] Disable logging requests.", deprecated=True, ) current_platform.pre_register_and_update(parser) return parser enable_log_requests class-attribute instance-attribute ¶ enable_log_requests: bool = False __init__ ¶ __init__( model: str = model, served_model_name: str | list[str] | None = served_model_name, tokenizer: str | None = tokenizer, hf_config_path: str | None = hf_config_path, runner: RunnerOption = runner, convert: ConvertOption = convert, skip_tokenizer_init: bool = skip_tokenizer_init, enable_prompt_embeds: bool = enable_prompt_embeds, tokenizer_mode: TokenizerMode | str = tokenizer_mode, trust_remote_code: bool = trust_remote_code, allowed_local_media_path: str = allowed_local_media_path, allowed_media_domains: list[str] | None = allowed_media_domains, download_dir: str | None = download_dir, safetensors_load_strategy: str = safetensors_load_strategy, load_format: str | LoadFormats = load_format, config_format: str = config_format, dtype: ModelDType = dtype, kv_cache_dtype: CacheDType = cache_dtype, seed: int = seed, max_model_len: int | None = max_model_len, cudagraph_capture_sizes: list[int] | None = cudagraph_capture_sizes, max_cudagraph_capture_size: int | None = get_field( CompilationConfig, "max_cudagraph_capture_size" ), distributed_executor_backend: str | DistributedExecutorBackend | type[Executor] | None = distributed_executor_backend, pipeline_parallel_size: int = pipeline_parallel_size, master_addr: str = master_addr, master_port: int = master_port, nnodes: int = nnodes, node_rank: int = node_rank, tensor_parallel_size: int = tensor_parallel_size, prefill_context_parallel_size: int = prefill_context_parallel_size, decode_context_parallel_size: int = decode_context_parallel_size, dcp_kv_cache_interleave_size: int = dcp_kv_cache_interleave_size, cp_kv_cache_interleave_size: int = cp_kv_cache_interleave_size, data_parallel_size: int = data_parallel_size, data_parallel_rank: int | None = None, data_parallel_start_rank: int | None = None, data_parallel_size_local: int | None = None, data_parallel_address: str | None = None, data_parallel_rpc_port: int | None = None, data_parallel_hybrid_lb: bool = False, data_parallel_external_lb: bool = False, data_parallel_backend: str = data_parallel_backend, enable_expert_parallel: bool = enable_expert_parallel, all2all_backend: str = all2all_backend, enable_dbo: bool = enable_dbo, ubatch_size: int = ubatch_size, dbo_decode_token_threshold: int = dbo_decode_token_threshold, dbo_prefill_token_threshold: int = dbo_prefill_token_threshold, disable_nccl_for_dp_synchronization: bool = disable_nccl_for_dp_synchronization, eplb_config: EPLBConfig = get_field( ParallelConfig, "eplb_config" ), enable_eplb: bool = enable_eplb, expert_placement_strategy: ExpertPlacementStrategy = expert_placement_strategy, _api_process_count: int = _api_process_count, _api_process_rank: int = _api_process_rank, max_parallel_loading_workers: int | None = max_parallel_loading_workers, block_size: BlockSize | None = block_size, enable_prefix_caching: bool | None = None, prefix_caching_hash_algo: PrefixCachingHashAlgo = prefix_caching_hash_algo, disable_sliding_window: bool = disable_sliding_window, disable_cascade_attn: bool = disable_cascade_attn, swap_space: float = swap_space, cpu_offload_gb: float = cpu_offload_gb, gpu_memory_utilization: float = gpu_memory_utilization, kv_cache_memory_bytes: int | None = kv_cache_memory_bytes, max_num_batched_tokens: int | None = None, max_num_partial_prefills: int = max_num_partial_prefills, max_long_partial_prefills: int = max_long_partial_prefills, long_prefill_token_threshold: int = long_prefill_token_threshold, max_num_seqs: int | None = None, max_logprobs: int = max_logprobs, logprobs_mode: LogprobsMode = logprobs_mode, disable_log_stats: bool = False, aggregate_engine_logging: bool = False, revision: str | None = revision, code_revision: str | None = code_revision, hf_token: bool | str | None = hf_token, hf_overrides: HfOverrides = get_field( ModelConfig, "hf_overrides" ), tokenizer_revision: str | None = tokenizer_revision, quantization: QuantizationMethods | None = quantization, enforce_eager: bool = enforce_eager, disable_custom_all_reduce: bool = disable_custom_all_reduce, limit_mm_per_prompt: dict[ str, int | dict[str, int] ] = get_field(MultiModalConfig, "limit_per_prompt"), enable_mm_embeds: bool = enable_mm_embeds, interleave_mm_strings: bool = interleave_mm_strings, media_io_kwargs: dict[str, dict[str, Any]] = get_field( MultiModalConfig, "media_io_kwargs" ), mm_processor_kwargs: dict[str, Any] | None = mm_processor_kwargs, mm_processor_cache_gb: float = mm_processor_cache_gb, mm_processor_cache_type: MMCacheType | None = mm_processor_cache_type, mm_shm_cache_max_object_size_mb: int = mm_shm_cache_max_object_size_mb, mm_encoder_tp_mode: MMEncoderTPMode = mm_encoder_tp_mode, mm_encoder_attn_backend: AttentionBackendEnum | str | None = mm_encoder_attn_backend, io_processor_plugin: str | None = None, skip_mm_profiling: bool = skip_mm_profiling, video_pruning_rate: float = video_pruning_rate, enable_lora: bool = False, max_loras: int = max_loras, max_lora_rank: int = max_lora_rank, default_mm_loras: dict[str, str] | None = default_mm_loras, fully_sharded_loras: bool = fully_sharded_loras, max_cpu_loras: int | None = max_cpu_loras, lora_dtype: str | dtype | None = lora_dtype, ray_workers_use_nsight: bool = ray_workers_use_nsight, num_gpu_blocks_override: int | None = num_gpu_blocks_override, model_loader_extra_config: dict = get_field( LoadConfig, "model_loader_extra_config" ), ignore_patterns: str | list[str] = get_field( LoadConfig, "ignore_patterns" ), enable_chunked_prefill: bool | None = None, disable_chunked_mm_input: bool = disable_chunked_mm_input, disable_hybrid_kv_cache_manager: bool | None = disable_hybrid_kv_cache_manager, structured_outputs_config: StructuredOutputsConfig = get_field( VllmConfig, "structured_outputs_config" ), reasoning_parser: str = reasoning_parser, reasoning_parser_plugin: str | None = None, logits_processor_pattern: str | None = logits_processor_pattern, speculative_config: dict[str, Any] | None = None, show_hidden_metrics_for_version: str | None = show_hidden_metrics_for_version, otlp_traces_endpoint: str | None = otlp_traces_endpoint, collect_detailed_traces: list[DetailedTraceModules] | None = collect_detailed_traces, kv_cache_metrics: bool = kv_cache_metrics, kv_cache_metrics_sample: float = get_field( ObservabilityConfig, "kv_cache_metrics_sample" ), cudagraph_metrics: bool = cudagraph_metrics, enable_layerwise_nvtx_tracing: bool = enable_layerwise_nvtx_tracing, enable_mfu_metrics: bool = enable_mfu_metrics, scheduling_policy: SchedulerPolicy = policy, scheduler_cls: str | type[object] | None = scheduler_cls, pooler_config: PoolerConfig | None = pooler_config, compilation_config: CompilationConfig = get_field( VllmConfig, "compilation_config" ), attention_config: AttentionConfig = get_field( VllmConfig, "attention_config" ), worker_cls: str = worker_cls, worker_extension_cls: str = worker_extension_cls, profiler_config: ProfilerConfig = get_field( VllmConfig, "profiler_config" ), kv_transfer_config: KVTransferConfig | None = None, kv_events_config: KVEventsConfig | None = None, ec_transfer_config: ECTransferConfig | None = None, generation_config: str = generation_config, enable_sleep_mode: bool = enable_sleep_mode, override_generation_config: dict[str, Any] = get_field( ModelConfig, "override_generation_config" ), model_impl: str = model_impl, override_attention_dtype: str = override_attention_dtype, attention_backend: AttentionBackendEnum | None = backend, calculate_kv_scales: bool = calculate_kv_scales, mamba_cache_dtype: MambaDType = mamba_cache_dtype, mamba_ssm_cache_dtype: MambaDType = mamba_ssm_cache_dtype, mamba_block_size: int | None = get_field( CacheConfig, "mamba_block_size" ), additional_config: dict[str, Any] = get_field( VllmConfig, "additional_config" ), use_tqdm_on_load: bool = use_tqdm_on_load, pt_load_map_location: str = pt_load_map_location, logits_processors: list[str | type[LogitsProcessor]] | None = logits_processors, async_scheduling: bool | None = async_scheduling, stream_interval: int = stream_interval, kv_sharing_fast_prefill: bool = kv_sharing_fast_prefill, optimization_level: OptimizationLevel = optimization_level, kv_offloading_size: float | None = kv_offloading_size, kv_offloading_backend: KVOffloadingBackend | None = kv_offloading_backend, tokens_only: bool = False, enable_log_requests: bool = False, ) -> None add_cli_args staticmethod ¶ add_cli_args( parser: FlexibleArgumentParser, async_args_only: bool = False, ) -> FlexibleArgumentParser Source code in vllm/engine/arg_utils.py 2008 2009 2010 2011 2012 2013 2014 2015 2016 2017 2018 2019 2020 2021 2022 2023 2024 2025 2026 2027 2028 2029 2030 2031 2032@staticmethod def add_cli_args( parser: FlexibleArgumentParser, async_args_only: bool = False ) -> FlexibleArgumentParser: # Initialize plugin to update the parser, for example, The plugin may # add a new kind of quantization method to --quantization argument or # a new device to --device argument. load_general_plugins() if not async_args_only: parser = EngineArgs.add_cli_args(parser) parser.add_argument( "--enable-log-requests", action=argparse.BooleanOptionalAction, default=AsyncEngineArgs.enable_log_requests, help="Enable logging requests.", ) parser.add_argument( "--disable-log-requests", action=argparse.BooleanOptionalAction, default=not AsyncEngineArgs.enable_log_requests, help="[DEPRECATED] Disable logging requests.", deprecated=True, ) current_platform.pre_register_and_update(parser) return parser EngineArgs dataclass ¶ Arguments for vLLM engine. Source code in vllm/engine/arg_utils.py 352 353 354 355 356 357 358 359 360 361 362 363 364 365 366 367 368 369 370 371 372 373 374 375 376 377 378 379 380 381 382 383 384 385 386 387 388 389 390 391 392 393 394 395 396 397 398 399 400 401 402 403 404 405 406 407 408 409 410 411 412 413 414 415 416 417 418 419 420 421 422 423 424 425 426 427 428 429 430 431 432 433 434 435 436 437 438 439 440 441 442 443 444 445 446 447 448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 465 466 467 468 469 470 471 472 473 474 475 476 477 478 479 480 481 482 483 484 485 486 487 488 489 490 491 492 493 494 495 496 497 498 499 500 501 502 503 504 505 506 507 508 509 510 511 512 513 514 515 516 517 518 519 520 521 522 523 524 525 526 527 528 529 530 531 532 533 534 535 536 537 538 539 540 541 542 543 544 545 546 547 548 549 550 551 552 553 554 555 556 557 558 559 560 561 562 563 564 565 566 567 568 569 570 571 572 573 574 575 576 577 578 579 580 581 582 583 584 585 586 587 588 589 590 591 592 593 594 595 596 597 598 599 600 601 602 603 604 605 606 607 608 609 610 611 612 613 614 615 616 617 618 619 620 621 622 623 624 625 626 627 628 629 630 631 632 633 634 635 636 637 638 639 640 641 642 643 644 645 646 647 648 649 650 651 652 653 654 655 656 657 658 659 660 661 662 663 664 665 666 667 668 669 670 671 672 673 674 675 676 677 678 679 680 681 682 683 684 685 686 687 688 689 690 691 692 693 694 695 696 697 698 699 700 701 702 703 704 705 706 707 708 709 710 711 712 713 714 715 716 717 718 719 720 721 722 723 724 725 726 727 728 729 730 731 732 733 734 735 736 737 738 739 740 741 742 743 744 745 746 747 748 749 750 751 752 753 754 755 756 757 758 759 760 761 762 763 764 765 766 767 768 769 770 771 772 773 774 775 776 777 778 779 780 781 782 783 784 785 786 787 788 789 790 791 792 793 794 795 796 797 798 799 800 801 802 803 804 805 806 807 808 809 810 811 812 813 814 815 816 817 818 819 820 821 822 823 824 825 826 827 828 829 830 831 832 833 834 835 836 837 838 839 840 841 842 843 844 845 846 847 848 849 850 851 852 853 854 855 856 857 858 859 860 861 862 863 864 865 866 867 868 869 870 871 872 873 874 875 876 877 878 879 880 881 882 883 884 885 886 887 888 889 890 891 892 893 894 895 896 897 898 899 900 901 902 903 904 905 906 907 908 909 910 911 912 913 914 915 916 917 918 919 920 921 922 923 924 925 926 927 928 929 930 931 932 933 934 935 936 937 938 939 940 941 942 943 944 945 946 947 948 949 950 951 952 953 954 955 956 957 958 959 960 961 962 963 964 965 966 967 968 969 970 971 972 973 974 975 976 977 978 979 980 981 982 983 984 985 986 987 988 989 990 991 992 993 994 995 996 997 998 999 1000 1001 1002 1003 1004 1005 1006 1007 1008 1009 1010 1011 1012 1013 1014 1015 1016 1017 1018 1019 1020 1021 1022 1023 1024 1025 1026 1027 1028 1029 1030 1031 1032 1033 1034 1035 1036 1037 1038 1039 1040 1041 1042 1043 1044 1045 1046 1047 1048 1049 1050 1051 1052 1053 1054 1055 1056 1057 1058 1059 1060 1061 1062 1063 1064 1065 1066 1067 1068 1069 1070 1071 1072 1073 1074 1075 1076 1077 1078 1079 1080 1081 1082 1083 1084 1085 1086 1087 1088 1089 1090 1091 1092 1093 1094 1095 1096 1097 1098 1099 1100 1101 1102 1103 1104 1105 1106 1107 1108 1109 1110 1111 1112 1113 1114 1115 1116 1117 1118 1119 1120 1121 1122 1123 1124 1125 1126 1127 1128 1129 1130 1131 1132 1133 1134 1135 1136 1137 1138 1139 1140 1141 1142 1143 1144 1145 1146 1147 1148 1149 1150 1151 1152 1153 1154 1155 1156 1157 1158 1159 1160 1161 1162 1163 1164 1165 1166 1167 1168 1169 1170 1171 1172 1173 1174 1175 1176 1177 1178 1179 1180 1181 1182 1183 1184 1185 1186 1187 1188 1189 1190 1191 1192 1193 1194 1195 1196 1197 1198 1199 1200 1201 1202 1203 1204 1205 1206 1207 1208 1209 1210 1211 1212 1213 1214 1215 1216 1217 1218 1219 1220 1221 1222 1223 1224 1225 1226 1227 1228 1229 1230 1231 1232 1233 1234 1235 1236 1237 1238 1239 1240 1241 1242 1243 1244 1245 1246 1247 1248 1249 1250 1251 1252 1253 1254 1255 1256 1257 1258 1259 1260 1261 1262 1263 1264 1265 1266 1267 1268 1269 1270 1271 1272 1273 1274 1275 1276 1277 1278 1279 1280 1281 1282 1283 1284 1285 1286 1287 1288 1289 1290 1291 1292 1293 1294 1295 1296 1297 1298 1299 1300 1301 1302 1303 1304 1305 1306 1307 1308 1309 1310 1311 1312 1313 1314 1315 1316 1317 1318 1319 1320 1321 1322 1323 1324 1325 1326 1327 1328 1329 1330 1331 1332 1333 1334 1335 1336 1337 1338 1339 1340 1341 1342 1343 1344 1345 1346 1347 1348 1349 1350 1351 1352 1353 1354 1355 1356 1357 1358 1359 1360 1361 1362 1363 1364 1365 1366 1367 1368 1369 1370 1371 1372 1373 1374 1375 1376 1377 1378 1379 1380 1381 1382 1383 1384 1385 1386 1387 1388 1389 1390 1391 1392 1393 1394 1395 1396 1397 1398 1399 1400 1401 1402 1403 1404 1405 1406 1407 1408 1409 1410 1411 1412 1413 1414 1415 1416 1417 1418 1419 1420 1421 1422 1423 1424 1425 1426 1427 1428 1429 1430 1431 1432 1433 1434 1435 1436 1437 1438 1439 1440 1441 1442 1443 1444 1445 1446 1447 1448 1449 1450 1451 1452 1453 1454 1455 1456 1457 1458 1459 1460 1461 1462 1463 1464 1465 1466 1467 1468 1469 1470 1471 1472 1473 1474 1475 1476 1477 1478 1479 1480 1481 1482 1483 1484 1485 1486 1487 1488 1489 1490 1491 1492 1493 1494 1495 1496 1497 1498 1499 1500 1501 1502 1503 1504 1505 1506 1507 1508 1509 1510 1511 1512 1513 1514 1515 1516 1517 1518 1519 1520 1521 1522 1523 1524 1525 1526 1527 1528 1529 1530 1531 1532 1533 1534 1535 1536 1537 1538 1539 1540 1541 1542 1543 1544 1545 1546 1547 1548 1549 1550 1551 1552 1553 1554 1555 1556 1557 1558 1559 1560 1561 1562 1563 1564 1565 1566 1567 1568 1569 1570 1571 1572 1573 1574 1575 1576 1577 1578 1579 1580 1581 1582 1583 1584 1585 1586 1587 1588 1589 1590 1591 1592 1593 1594 1595 1596 1597 1598 1599 1600 1601 1602 1603 1604 1605 1606 1607 1608 1609 1610 1611 1612 1613 1614 1615 1616 1617 1618 1619 1620 1621 1622 1623 1624 1625 1626 1627 1628 1629 1630 1631 1632 1633 1634 1635 1636 1637 1638 1639 1640 1641 1642 1643 1644 1645 1646 1647 1648 1649 1650 1651 1652 1653 1654 1655 1656 1657 1658 1659 1660 1661 1662 1663 1664 1665 1666 1667 1668 1669 1670 1671 1672 1673 1674 1675 1676 1677 1678 1679 1680 1681 1682 1683 1684 1685 1686 1687 1688 1689 1690 1691 1692 1693 1694 1695 1696 1697 1698 1699 1700 1701 1702 1703 1704 1705 1706 1707 1708 1709 1710 1711 1712 1713 1714 1715 1716 1717 1718 1719 1720 1721 1722 1723 1724 1725 1726 1727 1728 1729 1730 1731 1732 1733 1734 1735 1736 1737 1738 1739 1740 1741 1742 1743 1744 1745 1746 1747 1748 1749 1750 1751 1752 1753 1754 1755 1756 1757 1758 1759 1760 1761 1762 1763 1764 1765 1766 1767 1768 1769 1770 1771 1772 1773 1774 1775 1776 1777 1778 1779 1780 1781 1782 1783 1784 1785 1786 1787 1788 1789 1790 1791 1792 1793 1794 1795 1796 1797 1798 1799 1800 1801 1802 1803 1804 1805 1806 1807 1808 1809 1810 1811 1812 1813 1814 1815 1816 1817 1818 1819 1820 1821 1822 1823 1824 1825 1826 1827 1828 1829 1830 1831 1832 1833 1834 1835 1836 1837 1838 1839 1840 1841 1842 1843 1844 1845 1846 1847 1848 1849 1850 1851 1852 1853 1854 1855 1856 1857 1858 1859 1860 1861 1862 1863 1864 1865 1866 1867 1868 1869 1870 1871 1872 1873 1874 1875 1876 1877 1878 1879 1880 1881 1882 1883 1884 1885 1886 1887 1888 1889 1890 1891 1892 1893 1894 1895 1896 1897 1898 1899 1900 1901 1902 1903 1904 1905 1906 1907 1908 1909 1910 1911 1912 1913 1914 1915 1916 1917 1918 1919 1920 1921 1922 1923 1924 1925 1926 1927 1928 1929 1930 1931 1932 1933 1934 1935 1936 1937 1938 1939 1940 1941 1942 1943 1944 1945 1946 1947 1948 1949 1950 1951 1952 1953 1954 1955 1956 1957 1958 1959 1960 1961 1962 1963 1964 1965 1966 1967 1968 1969 1970 1971 1972 1973 1974 1975 1976 1977 1978 1979 1980 1981 1982 1983 1984 1985 1986 1987 1988 1989 1990 1991 1992 1993 1994 1995 1996 1997 1998 1999@dataclass class EngineArgs: """Arguments for vLLM engine.""" model: str = ModelConfig.model served_model_name: str | list[str] | None = ModelConfig.served_model_name tokenizer: str | None = ModelConfig.tokenizer hf_config_path: str | None = ModelConfig.hf_config_path runner: RunnerOption = ModelConfig.runner convert: ConvertOption = ModelConfig.convert skip_tokenizer_init: bool = ModelConfig.skip_tokenizer_init enable_prompt_embeds: bool = ModelConfig.enable_prompt_embeds tokenizer_mode: TokenizerMode | str = ModelConfig.tokenizer_mode trust_remote_code: bool = ModelConfig.trust_remote_code allowed_local_media_path: str = ModelConfig.allowed_local_media_path allowed_media_domains: list[str] | None = ModelConfig.allowed_media_domains download_dir: str | None = LoadConfig.download_dir safetensors_load_strategy: str = LoadConfig.safetensors_load_strategy load_format: str | LoadFormats = LoadConfig.load_format config_format: str = ModelConfig.config_format dtype: ModelDType = ModelConfig.dtype kv_cache_dtype: CacheDType = CacheConfig.cache_dtype seed: int = ModelConfig.seed max_model_len: int | None = ModelConfig.max_model_len cudagraph_capture_sizes: list[int] | None = ( CompilationConfig.cudagraph_capture_sizes ) max_cudagraph_capture_size: int | None = get_field( CompilationConfig, "max_cudagraph_capture_size" ) # Note: Specifying a custom executor backend by passing a class # is intended for expert use only. The API may change without # notice. distributed_executor_backend: ( str | DistributedExecutorBackend | type[Executor] | None ) = ParallelConfig.distributed_executor_backend # number of P/D disaggregation (or other disaggregation) workers pipeline_parallel_size: int = ParallelConfig.pipeline_parallel_size master_addr: str = ParallelConfig.master_addr master_port: int = ParallelConfig.master_port nnodes: int = ParallelConfig.nnodes node_rank: int = ParallelConfig.node_rank tensor_parallel_size: int = ParallelConfig.tensor_parallel_size prefill_context_parallel_size: int = ParallelConfig.prefill_context_parallel_size decode_context_parallel_size: int = ParallelConfig.decode_context_parallel_size dcp_kv_cache_interleave_size: int = ParallelConfig.dcp_kv_cache_interleave_size cp_kv_cache_interleave_size: int = ParallelConfig.cp_kv_cache_interleave_size data_parallel_size: int = ParallelConfig.data_parallel_size data_parallel_rank: int | None = None data_parallel_start_rank: int | None = None data_parallel_size_local: int | None = None data_parallel_address: str | None = None data_parallel_rpc_port: int | None = None data_parallel_hybrid_lb: bool = False data_parallel_external_lb: bool = False data_parallel_backend: str = ParallelConfig.data_parallel_backend enable_expert_parallel: bool = ParallelConfig.enable_expert_parallel all2all_backend: str = ParallelConfig.all2all_backend enable_dbo: bool = ParallelConfig.enable_dbo ubatch_size: int = ParallelConfig.ubatch_size dbo_decode_token_threshold: int = ParallelConfig.dbo_decode_token_threshold dbo_prefill_token_threshold: int = ParallelConfig.dbo_prefill_token_threshold disable_nccl_for_dp_synchronization: bool = ( ParallelConfig.disable_nccl_for_dp_synchronization ) eplb_config: EPLBConfig = get_field(ParallelConfig, "eplb_config") enable_eplb: bool = ParallelConfig.enable_eplb expert_placement_strategy: ExpertPlacementStrategy = ( ParallelConfig.expert_placement_strategy ) _api_process_count: int = ParallelConfig._api_process_count _api_process_rank: int = ParallelConfig._api_process_rank max_parallel_loading_workers: int | None = ( ParallelConfig.max_parallel_loading_workers ) block_size: BlockSize | None = CacheConfig.block_size enable_prefix_caching: bool | None = None prefix_caching_hash_algo: PrefixCachingHashAlgo = ( CacheConfig.prefix_caching_hash_algo ) disable_sliding_window: bool = ModelConfig.disable_sliding_window disable_cascade_attn: bool = ModelConfig.disable_cascade_attn swap_space: float = CacheConfig.swap_space cpu_offload_gb: float = CacheConfig.cpu_offload_gb gpu_memory_utilization: float = CacheConfig.gpu_memory_utilization kv_cache_memory_bytes: int | None = CacheConfig.kv_cache_memory_bytes max_num_batched_tokens: int | None = None max_num_partial_prefills: int = SchedulerConfig.max_num_partial_prefills max_long_partial_prefills: int = SchedulerConfig.max_long_partial_prefills long_prefill_token_threshold: int = SchedulerConfig.long_prefill_token_threshold max_num_seqs: int | None = None max_logprobs: int = ModelConfig.max_logprobs logprobs_mode: LogprobsMode = ModelConfig.logprobs_mode disable_log_stats: bool = False aggregate_engine_logging: bool = False revision: str | None = ModelConfig.revision code_revision: str | None = ModelConfig.code_revision hf_token: bool | str | None = ModelConfig.hf_token hf_overrides: HfOverrides = get_field(ModelConfig, "hf_overrides") tokenizer_revision: str | None = ModelConfig.tokenizer_revision quantization: QuantizationMethods | None = ModelConfig.quantization enforce_eager: bool = ModelConfig.enforce_eager disable_custom_all_reduce: bool = ParallelConfig.disable_custom_all_reduce limit_mm_per_prompt: dict[str, int | dict[str, int]] = get_field( MultiModalConfig, "limit_per_prompt" ) enable_mm_embeds: bool = MultiModalConfig.enable_mm_embeds interleave_mm_strings: bool = MultiModalConfig.interleave_mm_strings media_io_kwargs: dict[str, dict[str, Any]] = get_field( MultiModalConfig, "media_io_kwargs" ) mm_processor_kwargs: dict[str, Any] | None = MultiModalConfig.mm_processor_kwargs mm_processor_cache_gb: float = MultiModalConfig.mm_processor_cache_gb mm_processor_cache_type: MMCacheType | None = ( MultiModalConfig.mm_processor_cache_type ) mm_shm_cache_max_object_size_mb: int = ( MultiModalConfig.mm_shm_cache_max_object_size_mb ) mm_encoder_tp_mode: MMEncoderTPMode = MultiModalConfig.mm_encoder_tp_mode mm_encoder_attn_backend: AttentionBackendEnum | str | None = ( MultiModalConfig.mm_encoder_attn_backend ) io_processor_plugin: str | None = None skip_mm_profiling: bool = MultiModalConfig.skip_mm_profiling video_pruning_rate: float = MultiModalConfig.video_pruning_rate # LoRA fields enable_lora: bool = False max_loras: int = LoRAConfig.max_loras max_lora_rank: int = LoRAConfig.max_lora_rank default_mm_loras: dict[str, str] | None = LoRAConfig.default_mm_loras fully_sharded_loras: bool = LoRAConfig.fully_sharded_loras max_cpu_loras: int | None = LoRAConfig.max_cpu_loras lora_dtype: str | torch.dtype | None = LoRAConfig.lora_dtype ray_workers_use_nsight: bool = ParallelConfig.ray_workers_use_nsight num_gpu_blocks_override: int | None = CacheConfig.num_gpu_blocks_override model_loader_extra_config: dict = get_field(LoadConfig, "model_loader_extra_config") ignore_patterns: str | list[str] = get_field(LoadConfig, "ignore_patterns") enable_chunked_prefill: bool | None = None disable_chunked_mm_input: bool = SchedulerConfig.disable_chunked_mm_input disable_hybrid_kv_cache_manager: bool | None = ( SchedulerConfig.disable_hybrid_kv_cache_manager ) structured_outputs_config: StructuredOutputsConfig = get_field( VllmConfig, "structured_outputs_config" ) reasoning_parser: str = StructuredOutputsConfig.reasoning_parser reasoning_parser_plugin: str | None = None logits_processor_pattern: str | None = ModelConfig.logits_processor_pattern speculative_config: dict[str, Any] | None = None show_hidden_metrics_for_version: str | None = ( ObservabilityConfig.show_hidden_metrics_for_version ) otlp_traces_endpoint: str | None = ObservabilityConfig.otlp_traces_endpoint collect_detailed_traces: list[DetailedTraceModules] | None = ( ObservabilityConfig.collect_detailed_traces ) kv_cache_metrics: bool = ObservabilityConfig.kv_cache_metrics kv_cache_metrics_sample: float = get_field( ObservabilityConfig, "kv_cache_metrics_sample" ) cudagraph_metrics: bool = ObservabilityConfig.cudagraph_metrics enable_layerwise_nvtx_tracing: bool = ( ObservabilityConfig.enable_layerwise_nvtx_tracing ) enable_mfu_metrics: bool = ObservabilityConfig.enable_mfu_metrics scheduling_policy: SchedulerPolicy = SchedulerConfig.policy scheduler_cls: str | type[object] | None = SchedulerConfig.scheduler_cls pooler_config: PoolerConfig | None = ModelConfig.pooler_config compilation_config: CompilationConfig = get_field(VllmConfig, "compilation_config") attention_config: AttentionConfig = get_field(VllmConfig, "attention_config") worker_cls: str = ParallelConfig.worker_cls worker_extension_cls: str = ParallelConfig.worker_extension_cls profiler_config: ProfilerConfig = get_field(VllmConfig, "profiler_config") kv_transfer_config: KVTransferConfig | None = None kv_events_config: KVEventsConfig | None = None ec_transfer_config: ECTransferConfig | None = None generation_config: str = ModelConfig.generation_config enable_sleep_mode: bool = ModelConfig.enable_sleep_mode override_generation_config: dict[str, Any] = get_field( ModelConfig, "override_generation_config" ) model_impl: str = ModelConfig.model_impl override_attention_dtype: str = ModelConfig.override_attention_dtype attention_backend: AttentionBackendEnum | None = AttentionConfig.backend calculate_kv_scales: bool = CacheConfig.calculate_kv_scales mamba_cache_dtype: MambaDType = CacheConfig.mamba_cache_dtype mamba_ssm_cache_dtype: MambaDType = CacheConfig.mamba_ssm_cache_dtype mamba_block_size: int | None = get_field(CacheConfig, "mamba_block_size") additional_config: dict[str, Any] = get_field(VllmConfig, "additional_config") use_tqdm_on_load: bool = LoadConfig.use_tqdm_on_load pt_load_map_location: str = LoadConfig.pt_load_map_location logits_processors: list[str | type[LogitsProcessor]] | None = ( ModelConfig.logits_processors ) """Custom logitproc types""" async_scheduling: bool | None = SchedulerConfig.async_scheduling stream_interval: int = SchedulerConfig.stream_interval kv_sharing_fast_prefill: bool = CacheConfig.kv_sharing_fast_prefill optimization_level: OptimizationLevel = VllmConfig.optimization_level kv_offloading_size: float | None = CacheConfig.kv_offloading_size kv_offloading_backend: KVOffloadingBackend | None = ( CacheConfig.kv_offloading_backend ) tokens_only: bool = False def __post_init__(self): # support `EngineArgs(compilation_config={...})` # without having to manually construct a # CompilationConfig object if isinstance(self.compilation_config, dict): self.compilation_config = CompilationConfig(**self.compilation_config) if isinstance(self.attention_config, dict): self.attention_config = AttentionConfig(**self.attention_config) if isinstance(self.eplb_config, dict): self.eplb_config = EPLBConfig(**self.eplb_config) # Setup plugins from vllm.plugins import load_general_plugins load_general_plugins() # when use hf offline,replace model and tokenizer id to local model path if huggingface_hub.constants.HF_HUB_OFFLINE: model_id = self.model self.model = get_model_path(self.model, self.revision) if model_id is not self.model: logger.info( "HF_HUB_OFFLINE is True, replace model_id [%s] to model_path [%s]", model_id, self.model, ) if self.tokenizer is not None: tokenizer_id = self.tokenizer self.tokenizer = get_model_path(self.tokenizer, self.tokenizer_revision) if tokenizer_id is not self.tokenizer: logger.info( "HF_HUB_OFFLINE is True, replace tokenizer_id [%s] " "to tokenizer_path [%s]", tokenizer_id, self.tokenizer, ) @staticmethod def add_cli_args(parser: FlexibleArgumentParser) -> FlexibleArgumentParser: """Shared CLI arguments for vLLM engine.""" # Model arguments model_kwargs = get_kwargs(ModelConfig) model_group = parser.add_argument_group( title="ModelConfig", description=ModelConfig.__doc__, ) if not ("serve" in sys.argv[1:] and "--help" in sys.argv[1:]): model_group.add_argument("--model", **model_kwargs["model"]) model_group.add_argument("--runner", **model_kwargs["runner"]) model_group.add_argument("--convert", **model_kwargs["convert"]) model_group.add_argument("--tokenizer", **model_kwargs["tokenizer"]) model_group.add_argument("--tokenizer-mode", **model_kwargs["tokenizer_mode"]) model_group.add_argument( "--trust-remote-code", **model_kwargs["trust_remote_code"] ) model_group.add_argument("--dtype", **model_kwargs["dtype"]) model_group.add_argument("--seed", **model_kwargs["seed"]) model_group.add_argument("--hf-config-path", **model_kwargs["hf_config_path"]) model_group.add_argument( "--allowed-local-media-path", **model_kwargs["allowed_local_media_path"] ) model_group.add_argument( "--allowed-media-domains", **model_kwargs["allowed_media_domains"] ) model_group.add_argument("--revision", **model_kwargs["revision"]) model_group.add_argument("--code-revision", **model_kwargs["code_revision"]) model_group.add_argument( "--tokenizer-revision", **model_kwargs["tokenizer_revision"] ) model_group.add_argument("--max-model-len", **model_kwargs["max_model_len"]) model_group.add_argument("--quantization", "-q", **model_kwargs["quantization"]) model_group.add_argument("--enforce-eager", **model_kwargs["enforce_eager"]) model_group.add_argument("--max-logprobs", **model_kwargs["max_logprobs"]) model_group.add_argument("--logprobs-mode", **model_kwargs["logprobs_mode"]) model_group.add_argument( "--disable-sliding-window", **model_kwargs["disable_sliding_window"] ) model_group.add_argument( "--disable-cascade-attn", **model_kwargs["disable_cascade_attn"] ) model_group.add_argument( "--skip-tokenizer-init", **model_kwargs["skip_tokenizer_init"] ) model_group.add_argument( "--enable-prompt-embeds", **model_kwargs["enable_prompt_embeds"] ) model_group.add_argument( "--served-model-name", **model_kwargs["served_model_name"] ) model_group.add_argument("--config-format", **model_kwargs["config_format"]) # This one is a special case because it can bool # or str. TODO: Handle this in get_kwargs model_group.add_argument( "--hf-token", type=str, nargs="?", const=True, default=model_kwargs["hf_token"]["default"], help=model_kwargs["hf_token"]["help"], ) model_group.add_argument("--hf-overrides", **model_kwargs["hf_overrides"]) model_group.add_argument("--pooler-config", **model_kwargs["pooler_config"]) model_group.add_argument( "--logits-processor-pattern", **model_kwargs["logits_processor_pattern"] ) model_group.add_argument( "--generation-config", **model_kwargs["generation_config"] ) model_group.add_argument( "--override-generation-config", **model_kwargs["override_generation_config"] ) model_group.add_argument( "--enable-sleep-mode", **model_kwargs["enable_sleep_mode"] ) model_group.add_argument("--model-impl", **model_kwargs["model_impl"]) model_group.add_argument( "--override-attention-dtype", **model_kwargs["override_attention_dtype"] ) model_group.add_argument( "--logits-processors", **model_kwargs["logits_processors"] ) model_group.add_argument( "--io-processor-plugin", **model_kwargs["io_processor_plugin"] ) # Model loading arguments load_kwargs = get_kwargs(LoadConfig) load_group = parser.add_argument_group( title="LoadConfig", description=LoadConfig.__doc__, ) load_group.add_argument("--load-format", **load_kwargs["load_format"]) load_group.add_argument("--download-dir", **load_kwargs["download_dir"]) load_group.add_argument( "--safetensors-load-strategy", **load_kwargs["safetensors_load_strategy"] ) load_group.add_argument( "--model-loader-extra-config", **load_kwargs["model_loader_extra_config"] ) load_group.add_argument("--ignore-patterns", **load_kwargs["ignore_patterns"]) load_group.add_argument("--use-tqdm-on-load", **load_kwargs["use_tqdm_on_load"]) load_group.add_argument( "--pt-load-map-location", **load_kwargs["pt_load_map_location"] ) # Attention arguments attention_kwargs = get_kwargs(AttentionConfig) attention_group = parser.add_argument_group( title="AttentionConfig", description=AttentionConfig.__doc__, ) attention_group.add_argument( "--attention-backend", **attention_kwargs["backend"] ) # Structured outputs arguments structured_outputs_kwargs = get_kwargs(StructuredOutputsConfig) structured_outputs_group = parser.add_argument_group( title="StructuredOutputsConfig", description=StructuredOutputsConfig.__doc__, ) structured_outputs_group.add_argument( "--reasoning-parser", # Choices need to be validated after parsing to include plugins **structured_outputs_kwargs["reasoning_parser"], ) structured_outputs_group.add_argument( "--reasoning-parser-plugin", **structured_outputs_kwargs["reasoning_parser_plugin"], ) # Parallel arguments parallel_kwargs = get_kwargs(ParallelConfig) parallel_group = parser.add_argument_group( title="ParallelConfig", description=ParallelConfig.__doc__, ) parallel_group.add_argument( "--distributed-executor-backend", **parallel_kwargs["distributed_executor_backend"], ) parallel_group.add_argument( "--pipeline-parallel-size", "-pp", **parallel_kwargs["pipeline_parallel_size"], ) parallel_group.add_argument("--master-addr", **parallel_kwargs["master_addr"]) parallel_group.add_argument("--master-port", **parallel_kwargs["master_port"]) parallel_group.add_argument("--nnodes", "-n", **parallel_kwargs["nnodes"]) parallel_group.add_argument("--node-rank", "-r", **parallel_kwargs["node_rank"]) parallel_group.add_argument( "--tensor-parallel-size", "-tp", **parallel_kwargs["tensor_parallel_size"] ) parallel_group.add_argument( "--decode-context-parallel-size", "-dcp", **parallel_kwargs["decode_context_parallel_size"], ) parallel_group.add_argument( "--dcp-kv-cache-interleave-size", **parallel_kwargs["dcp_kv_cache_interleave_size"], ) parallel_group.add_argument( "--cp-kv-cache-interleave-size", **parallel_kwargs["cp_kv_cache_interleave_size"], ) parallel_group.add_argument( "--prefill-context-parallel-size", "-pcp", **parallel_kwargs["prefill_context_parallel_size"], ) parallel_group.add_argument( "--data-parallel-size", "-dp", **parallel_kwargs["data_parallel_size"] ) parallel_group.add_argument( "--data-parallel-rank", "-dpn", type=int, help="Data parallel rank of this instance. " "When set, enables external load balancer mode.", ) parallel_group.add_argument( "--data-parallel-start-rank", "-dpr", type=int, help="Starting data parallel rank for secondary nodes.", ) parallel_group.add_argument( "--data-parallel-size-local", "-dpl", type=int, help="Number of data parallel replicas to run on this node.", ) parallel_group.add_argument( "--data-parallel-address", "-dpa", type=str, help="Address of data parallel cluster head-node.", ) parallel_group.add_argument( "--data-parallel-rpc-port", "-dpp", type=int, help="Port for data parallel RPC communication.", ) parallel_group.add_argument( "--data-parallel-backend", "-dpb", type=str, default="mp", help='Backend for data parallel, either "mp" or "ray".', ) parallel_group.add_argument( "--data-parallel-hybrid-lb", "-dph", **parallel_kwargs["data_parallel_hybrid_lb"], ) parallel_group.add_argument( "--data-parallel-external-lb", "-dpe", **parallel_kwargs["data_parallel_external_lb"], ) parallel_group.add_argument( "--enable-expert-parallel", **parallel_kwargs["enable_expert_parallel"] ) parallel_group.add_argument( "--all2all-backend", **parallel_kwargs["all2all_backend"] ) parallel_group.add_argument("--enable-dbo", **parallel_kwargs["enable_dbo"]) parallel_group.add_argument( "--ubatch-size", **parallel_kwargs["ubatch_size"], ) parallel_group.add_argument( "--dbo-decode-token-threshold", **parallel_kwargs["dbo_decode_token_threshold"], ) parallel_group.add_argument( "--dbo-prefill-token-threshold", **parallel_kwargs["dbo_prefill_token_threshold"], ) parallel_group.add_argument( "--disable-nccl-for-dp-synchronization", **parallel_kwargs["disable_nccl_for_dp_synchronization"], ) parallel_group.add_argument("--enable-eplb", **parallel_kwargs["enable_eplb"]) parallel_group.add_argument("--eplb-config", **parallel_kwargs["eplb_config"]) parallel_group.add_argument( "--expert-placement-strategy", **parallel_kwargs["expert_placement_strategy"], ) parallel_group.add_argument( "--max-parallel-loading-workers", **parallel_kwargs["max_parallel_loading_workers"], ) parallel_group.add_argument( "--ray-workers-use-nsight", **parallel_kwargs["ray_workers_use_nsight"] ) parallel_group.add_argument( "--disable-custom-all-reduce", **parallel_kwargs["disable_custom_all_reduce"], ) parallel_group.add_argument("--worker-cls", **parallel_kwargs["worker_cls"]) parallel_group.add_argument( "--worker-extension-cls", **parallel_kwargs["worker_extension_cls"] ) # KV cache arguments cache_kwargs = get_kwargs(CacheConfig) cache_group = parser.add_argument_group( title="CacheConfig", description=CacheConfig.__doc__, ) cache_group.add_argument("--block-size", **cache_kwargs["block_size"]) cache_group.add_argument( "--gpu-memory-utilization", **cache_kwargs["gpu_memory_utilization"] ) cache_group.add_argument( "--kv-cache-memory-bytes", **cache_kwargs["kv_cache_memory_bytes"] ) cache_group.add_argument("--swap-space", **cache_kwargs["swap_space"]) cache_group.add_argument("--kv-cache-dtype", **cache_kwargs["cache_dtype"]) cache_group.add_argument( "--num-gpu-blocks-override", **cache_kwargs["num_gpu_blocks_override"] ) cache_group.add_argument( "--enable-prefix-caching", **{ **cache_kwargs["enable_prefix_caching"], "default": None, }, ) cache_group.add_argument( "--prefix-caching-hash-algo", **cache_kwargs["prefix_caching_hash_algo"] ) cache_group.add_argument("--cpu-offload-gb", **cache_kwargs["cpu_offload_gb"]) cache_group.add_argument( "--calculate-kv-scales", **cache_kwargs["calculate_kv_scales"] ) cache_group.add_argument( "--kv-sharing-fast-prefill", **cache_kwargs["kv_sharing_fast_prefill"] ) cache_group.add_argument( "--mamba-cache-dtype", **cache_kwargs["mamba_cache_dtype"] ) cache_group.add_argument( "--mamba-ssm-cache-dtype", **cache_kwargs["mamba_ssm_cache_dtype"] ) cache_group.add_argument( "--mamba-block-size", **cache_kwargs["mamba_block_size"] ) cache_group.add_argument( "--kv-offloading-size", **cache_kwargs["kv_offloading_size"] ) cache_group.add_argument( "--kv-offloading-backend", **cache_kwargs["kv_offloading_backend"] ) # Multimodal related configs multimodal_kwargs = get_kwargs(MultiModalConfig) multimodal_group = parser.add_argument_group( title="MultiModalConfig", description=MultiModalConfig.__doc__, ) multimodal_group.add_argument( "--limit-mm-per-prompt", **multimodal_kwargs["limit_per_prompt"] ) multimodal_group.add_argument( "--enable-mm-embeds", **multimodal_kwargs["enable_mm_embeds"] ) multimodal_group.add_argument( "--media-io-kwargs", **multimodal_kwargs["media_io_kwargs"] ) multimodal_group.add_argument( "--mm-processor-kwargs", **multimodal_kwargs["mm_processor_kwargs"] ) multimodal_group.add_argument( "--mm-processor-cache-gb", **multimodal_kwargs["mm_processor_cache_gb"] ) multimodal_group.add_argument( "--mm-processor-cache-type", **multimodal_kwargs["mm_processor_cache_type"] ) multimodal_group.add_argument( "--mm-shm-cache-max-object-size-mb", **multimodal_kwargs["mm_shm_cache_max_object_size_mb"], ) multimodal_group.add_argument( "--mm-encoder-tp-mode", **multimodal_kwargs["mm_encoder_tp_mode"] ) multimodal_group.add_argument( "--mm-encoder-attn-backend", **multimodal_kwargs["mm_encoder_attn_backend"], ) multimodal_group.add_argument( "--interleave-mm-strings", **multimodal_kwargs["interleave_mm_strings"] ) multimodal_group.add_argument( "--skip-mm-profiling", **multimodal_kwargs["skip_mm_profiling"] ) multimodal_group.add_argument( "--video-pruning-rate", **multimodal_kwargs["video_pruning_rate"] ) # LoRA related configs lora_kwargs = get_kwargs(LoRAConfig) lora_group = parser.add_argument_group( title="LoRAConfig", description=LoRAConfig.__doc__, ) lora_group.add_argument( "--enable-lora", action=argparse.BooleanOptionalAction, help="If True, enable handling of LoRA adapters.", ) lora_group.add_argument("--max-loras", **lora_kwargs["max_loras"]) lora_group.add_argument("--max-lora-rank", **lora_kwargs["max_lora_rank"]) lora_group.add_argument( "--lora-dtype", **lora_kwargs["lora_dtype"], ) lora_group.add_argument("--max-cpu-loras", **lora_kwargs["max_cpu_loras"]) lora_group.add_argument( "--fully-sharded-loras", **lora_kwargs["fully_sharded_loras"] ) lora_group.add_argument("--default-mm-loras", **lora_kwargs["default_mm_loras"]) # Observability arguments observability_kwargs = get_kwargs(ObservabilityConfig) observability_group = parser.add_argument_group( title="ObservabilityConfig", description=ObservabilityConfig.__doc__, ) observability_group.add_argument( "--show-hidden-metrics-for-version", **observability_kwargs["show_hidden_metrics_for_version"], ) observability_group.add_argument( "--otlp-traces-endpoint", **observability_kwargs["otlp_traces_endpoint"] ) # TODO: generalise this special case choices = observability_kwargs["collect_detailed_traces"]["choices"] metavar = f"{{{','.join(choices)}}}" observability_kwargs["collect_detailed_traces"]["metavar"] = metavar observability_kwargs["collect_detailed_traces"]["choices"] += [ ",".join(p) for p in permutations(get_args(DetailedTraceModules), r=2) ] observability_group.add_argument( "--collect-detailed-traces", **observability_kwargs["collect_detailed_traces"], ) observability_group.add_argument( "--kv-cache-metrics", **observability_kwargs["kv_cache_metrics"] ) observability_group.add_argument( "--kv-cache-metrics-sample", **observability_kwargs["kv_cache_metrics_sample"], ) observability_group.add_argument( "--cudagraph-metrics", **observability_kwargs["cudagraph_metrics"], ) observability_group.add_argument( "--enable-layerwise-nvtx-tracing", **observability_kwargs["enable_layerwise_nvtx_tracing"], ) observability_group.add_argument( "--enable-mfu-metrics", **observability_kwargs["enable_mfu_metrics"], ) # Scheduler arguments scheduler_kwargs = get_kwargs(SchedulerConfig) scheduler_group = parser.add_argument_group( title="SchedulerConfig", description=SchedulerConfig.__doc__, ) scheduler_group.add_argument( "--max-num-batched-tokens", **{ **scheduler_kwargs["max_num_batched_tokens"], "default": None, }, ) scheduler_group.add_argument( "--max-num-seqs", **{ **scheduler_kwargs["max_num_seqs"], "default": None, }, ) scheduler_group.add_argument( "--max-num-partial-prefills", **scheduler_kwargs["max_num_partial_prefills"] ) scheduler_group.add_argument( "--max-long-partial-prefills", **scheduler_kwargs["max_long_partial_prefills"], ) scheduler_group.add_argument( "--long-prefill-token-threshold", **scheduler_kwargs["long_prefill_token_threshold"], ) # multi-step scheduling has been removed; corresponding arguments # are no longer supported. scheduler_group.add_argument( "--scheduling-policy", **scheduler_kwargs["policy"] ) scheduler_group.add_argument( "--enable-chunked-prefill", **{ **scheduler_kwargs["enable_chunked_prefill"], "default": None, }, ) scheduler_group.add_argument( "--disable-chunked-mm-input", **scheduler_kwargs["disable_chunked_mm_input"] ) scheduler_group.add_argument( "--scheduler-cls", **scheduler_kwargs["scheduler_cls"] ) scheduler_group.add_argument( "--disable-hybrid-kv-cache-manager", **scheduler_kwargs["disable_hybrid_kv_cache_manager"], ) scheduler_group.add_argument( "--async-scheduling", **scheduler_kwargs["async_scheduling"] ) scheduler_group.add_argument( "--stream-interval", **scheduler_kwargs["stream_interval"] ) # Compilation arguments compilation_kwargs = get_kwargs(CompilationConfig) compilation_group = parser.add_argument_group( title="CompilationConfig", description=CompilationConfig.__doc__, ) compilation_group.add_argument( "--cudagraph-capture-sizes", **compilation_kwargs["cudagraph_capture_sizes"] ) compilation_group.add_argument( "--max-cudagraph-capture-size", **compilation_kwargs["max_cudagraph_capture_size"], ) # vLLM arguments vllm_kwargs = get_kwargs(VllmConfig) vllm_group = parser.add_argument_group( title="VllmConfig", description=VllmConfig.__doc__, ) # We construct SpeculativeConfig using fields from other configs in # create_engine_config. So we set the type to a JSON string here to # delay the Pydantic validation that comes with SpeculativeConfig. vllm_kwargs["speculative_config"]["type"] = optional_type(json.loads) vllm_group.add_argument( "--speculative-config", **vllm_kwargs["speculative_config"] ) vllm_group.add_argument( "--kv-transfer-config", **vllm_kwargs["kv_transfer_config"] ) vllm_group.add_argument("--kv-events-config", **vllm_kwargs["kv_events_config"]) vllm_group.add_argument( "--ec-transfer-config", **vllm_kwargs["ec_transfer_config"] ) vllm_group.add_argument( "--compilation-config", "-cc", **vllm_kwargs["compilation_config"] ) vllm_group.add_argument( "--attention-config", "-ac", **vllm_kwargs["attention_config"] ) vllm_group.add_argument( "--additional-config", **vllm_kwargs["additional_config"] ) vllm_group.add_argument( "--structured-outputs-config", **vllm_kwargs["structured_outputs_config"] ) vllm_group.add_argument("--profiler-config", **vllm_kwargs["profiler_config"]) vllm_group.add_argument( "--optimization-level", **vllm_kwargs["optimization_level"] ) # Other arguments parser.add_argument( "--disable-log-stats", action="store_true", help="Disable logging statistics.", ) parser.add_argument( "--aggregate-engine-logging", action="store_true", help="Log aggregate rather than per-engine statistics " "when using data parallelism.", ) return parser @classmethod def from_cli_args(cls, args: argparse.Namespace): # Get the list of attributes of this dataclass. attrs = [attr.name for attr in dataclasses.fields(cls)] # Set the attributes from the parsed arguments. engine_args = cls( **{attr: getattr(args, attr) for attr in attrs if hasattr(args, attr)} ) return engine_args def create_model_config(self) -> ModelConfig: # gguf file needs a specific model loader if is_gguf(self.model): self.quantization = self.load_format = "gguf" if not envs.VLLM_ENABLE_V1_MULTIPROCESSING: logger.warning( "The global random seed is set to %d. Since " "VLLM_ENABLE_V1_MULTIPROCESSING is set to False, this may " "affect the random state of the Python process that " "launched vLLM.", self.seed, ) return ModelConfig( model=self.model, hf_config_path=self.hf_config_path, runner=self.runner, convert=self.convert, tokenizer=self.tokenizer, tokenizer_mode=self.tokenizer_mode, trust_remote_code=self.trust_remote_code, allowed_local_media_path=self.allowed_local_media_path, allowed_media_domains=self.allowed_media_domains, dtype=self.dtype, seed=self.seed, revision=self.revision, code_revision=self.code_revision, hf_token=self.hf_token, hf_overrides=self.hf_overrides, tokenizer_revision=self.tokenizer_revision, max_model_len=self.max_model_len, quantization=self.quantization, enforce_eager=self.enforce_eager, max_logprobs=self.max_logprobs, logprobs_mode=self.logprobs_mode, disable_sliding_window=self.disable_sliding_window, disable_cascade_attn=self.disable_cascade_attn, skip_tokenizer_init=self.skip_tokenizer_init, enable_prompt_embeds=self.enable_prompt_embeds, served_model_name=self.served_model_name, limit_mm_per_prompt=self.limit_mm_per_prompt, enable_mm_embeds=self.enable_mm_embeds, interleave_mm_strings=self.interleave_mm_strings, media_io_kwargs=self.media_io_kwargs, skip_mm_profiling=self.skip_mm_profiling, config_format=self.config_format, mm_processor_kwargs=self.mm_processor_kwargs, mm_processor_cache_gb=self.mm_processor_cache_gb, mm_processor_cache_type=self.mm_processor_cache_type, mm_shm_cache_max_object_size_mb=self.mm_shm_cache_max_object_size_mb, mm_encoder_tp_mode=self.mm_encoder_tp_mode, mm_encoder_attn_backend=self.mm_encoder_attn_backend, pooler_config=self.pooler_config, logits_processor_pattern=self.logits_processor_pattern, generation_config=self.generation_config, override_generation_config=self.override_generation_config, enable_sleep_mode=self.enable_sleep_mode, model_impl=self.model_impl, override_attention_dtype=self.override_attention_dtype, logits_processors=self.logits_processors, video_pruning_rate=self.video_pruning_rate, io_processor_plugin=self.io_processor_plugin, ) def validate_tensorizer_args(self): from vllm.model_executor.model_loader.tensorizer import TensorizerConfig for key in self.model_loader_extra_config: if key in TensorizerConfig._fields: self.model_loader_extra_config["tensorizer_config"][key] = ( self.model_loader_extra_config[key] ) def create_load_config(self) -> LoadConfig: if self.quantization == "bitsandbytes": self.load_format = "bitsandbytes" if self.load_format == "tensorizer": if hasattr(self.model_loader_extra_config, "to_serializable"): self.model_loader_extra_config = ( self.model_loader_extra_config.to_serializable() ) self.model_loader_extra_config["tensorizer_config"] = {} self.model_loader_extra_config["tensorizer_config"]["tensorizer_dir"] = ( self.model ) self.validate_tensorizer_args() return LoadConfig( load_format=self.load_format, download_dir=self.download_dir, safetensors_load_strategy=self.safetensors_load_strategy, device="cpu" if is_online_quantization(self.quantization) else None, model_loader_extra_config=self.model_loader_extra_config, ignore_patterns=self.ignore_patterns, use_tqdm_on_load=self.use_tqdm_on_load, pt_load_map_location=self.pt_load_map_location, ) def create_speculative_config( self, target_model_config: ModelConfig, target_parallel_config: ParallelConfig, ) -> SpeculativeConfig | None: """Initializes and returns a SpeculativeConfig object based on `speculative_config`. This function utilizes `speculative_config` to create a SpeculativeConfig object. The `speculative_config` can either be provided as a JSON string input via CLI arguments or directly as a dictionary from the engine. """ if self.speculative_config is None: return None # Note(Shangming): These parameters are not obtained from the cli arg # '--speculative-config' and must be passed in when creating the engine # config. self.speculative_config.update( { "target_model_config": target_model_config, "target_parallel_config": target_parallel_config, } ) return SpeculativeConfig(**self.speculative_config) def create_engine_config( self, usage_context: UsageContext | None = None, headless: bool = False, ) -> VllmConfig: """ Create the VllmConfig. NOTE: If VllmConfig is incompatible, we raise an error. """ current_platform.pre_register_and_update() device_config = DeviceConfig(device=cast(Device, current_platform.device_type)) # Check if the model is a speculator and override model/tokenizer/config # BEFORE creating ModelConfig, so the config is created with the target model # Skip speculator detection for cloud storage models (eg: S3, GCS) since # HuggingFace cannot load configs directly from S3 URLs. S3 models can still # use speculators with explicit --speculative-config. if not is_cloud_storage(self.model): (self.model, self.tokenizer, self.speculative_config) = ( maybe_override_with_speculators( model=self.model, tokenizer=self.tokenizer, revision=self.revision, trust_remote_code=self.trust_remote_code, vllm_speculative_config=self.speculative_config, ) ) model_config = self.create_model_config() self.model = model_config.model self.tokenizer = model_config.tokenizer self._check_feature_supported(model_config) self._set_default_chunked_prefill_and_prefix_caching_args(model_config) self._set_default_max_num_seqs_and_batched_tokens_args( usage_context, model_config ) sliding_window: int | None = None if not is_interleaved(model_config.hf_text_config): # Only set CacheConfig.sliding_window if the model is all sliding # window. Otherwise CacheConfig.sliding_window will override the # global layers in interleaved sliding window models. sliding_window = model_config.get_sliding_window() # Note(hc): In the current implementation of decode context # parallel(DCP), tp_size needs to be divisible by dcp_size, # because the world size does not change by dcp, it simply # reuses the GPUs of TP group, and split one TP group into # tp_size//dcp_size DCP groups. assert self.tensor_parallel_size % self.decode_context_parallel_size == 0, ( f"tp_size={self.tensor_parallel_size} must be divisible by" f"dcp_size={self.decode_context_parallel_size}." ) # Resolve "auto" kv_cache_dtype to actual value from model config resolved_cache_dtype = resolve_kv_cache_dtype_string( self.kv_cache_dtype, model_config ) cache_config = CacheConfig( block_size=self.block_size, gpu_memory_utilization=self.gpu_memory_utilization, kv_cache_memory_bytes=self.kv_cache_memory_bytes, swap_space=self.swap_space, cache_dtype=resolved_cache_dtype, is_attention_free=model_config.is_attention_free, num_gpu_blocks_override=self.num_gpu_blocks_override, sliding_window=sliding_window, enable_prefix_caching=self.enable_prefix_caching, prefix_caching_hash_algo=self.prefix_caching_hash_algo, cpu_offload_gb=self.cpu_offload_gb, calculate_kv_scales=self.calculate_kv_scales, kv_sharing_fast_prefill=self.kv_sharing_fast_prefill, mamba_cache_dtype=self.mamba_cache_dtype, mamba_ssm_cache_dtype=self.mamba_ssm_cache_dtype, mamba_block_size=self.mamba_block_size, kv_offloading_size=self.kv_offloading_size, kv_offloading_backend=self.kv_offloading_backend, ) ray_runtime_env = None if is_ray_initialized(): # Ray Serve LLM calls `create_engine_config` in the context # of a Ray task, therefore we check is_ray_initialized() # as opposed to is_in_ray_actor(). import ray ray_runtime_env = ray.get_runtime_context().runtime_env # Avoid logging sensitive environment variables sanitized_env = ray_runtime_env.to_dict() if ray_runtime_env else {} if "env_vars" in sanitized_env: sanitized_env["env_vars"] = { k: "***" for k in sanitized_env["env_vars"] } logger.info("Using ray runtime env (env vars redacted): %s", sanitized_env) # Get the current placement group if Ray is initialized and # we are in a Ray actor. If so, then the placement group will be # passed to spawned processes. placement_group = None if is_in_ray_actor(): import ray # This call initializes Ray automatically if it is not initialized, # but we should not do this here. placement_group = ray.util.get_current_placement_group() assert not headless or not self.data_parallel_hybrid_lb, ( "data_parallel_hybrid_lb is not applicable in headless mode" ) assert not (self.data_parallel_hybrid_lb and self.data_parallel_external_lb), ( "data_parallel_hybrid_lb and data_parallel_external_lb cannot both be True." ) assert self.data_parallel_backend == "mp" or self.nnodes == 1, ( "nnodes > 1 is only supported with data_parallel_backend=mp" ) inferred_data_parallel_rank = 0 if self.nnodes > 1: world_size = ( self.data_parallel_size * self.pipeline_parallel_size * self.tensor_parallel_size ) world_size_within_dp = ( self.pipeline_parallel_size * self.tensor_parallel_size ) local_world_size = world_size // self.nnodes assert world_size % self.nnodes == 0, ( f"world_size={world_size} must be divisible by nnodes={self.nnodes}." ) assert self.node_rank < self.nnodes, ( f"node_rank={self.node_rank} must be less than nnodes={self.nnodes}." ) inferred_data_parallel_rank = ( self.node_rank * local_world_size ) // world_size_within_dp if self.data_parallel_size > 1 and self.data_parallel_external_lb: self.data_parallel_rank = inferred_data_parallel_rank logger.info( "Inferred data_parallel_rank %d from node_rank %d for external lb", self.data_parallel_rank, self.node_rank, ) elif self.data_parallel_size_local is None: # Infer data parallel size local for internal dplb: self.data_parallel_size_local = max( local_world_size // world_size_within_dp, 1 ) data_parallel_external_lb = ( self.data_parallel_external_lb or self.data_parallel_rank is not None ) # Local DP rank = 1, use pure-external LB. if data_parallel_external_lb: assert self.data_parallel_rank is not None, ( "data_parallel_rank or node_rank must be specified if " "data_parallel_external_lb is enable." ) assert self.data_parallel_size_local in (1, None), ( "data_parallel_size_local must be 1 or None when data_parallel_rank " "is set" ) data_parallel_size_local = 1 # Use full external lb if we have local_size of 1. self.data_parallel_hybrid_lb = False elif self.data_parallel_size_local is not None: data_parallel_size_local = self.data_parallel_size_local if self.data_parallel_start_rank and not headless: # Infer hybrid LB mode. self.data_parallel_hybrid_lb = True if self.data_parallel_hybrid_lb and data_parallel_size_local == 1: # Use full external lb if we have local_size of 1. logger.warning( "data_parallel_hybrid_lb is not eligible when " "data_parallel_size_local = 1, autoswitch to " "data_parallel_external_lb." ) data_parallel_external_lb = True self.data_parallel_hybrid_lb = False if data_parallel_size_local == self.data_parallel_size: # Disable hybrid LB mode if set for a single node self.data_parallel_hybrid_lb = False self.data_parallel_rank = ( self.data_parallel_start_rank or inferred_data_parallel_rank ) if self.nnodes > 1: logger.info( "Inferred data_parallel_rank %d from node_rank %d", self.data_parallel_rank, self.node_rank, ) else: assert not self.data_parallel_hybrid_lb, ( "data_parallel_size_local must be set to use data_parallel_hybrid_lb." ) if self.data_parallel_backend == "ray" and ( envs.VLLM_RAY_DP_PACK_STRATEGY == "span" ): # Data parallel size defaults to 1 if DP ranks are spanning # multiple nodes data_parallel_size_local = 1 else: # Otherwise local DP size defaults to global DP size if not set data_parallel_size_local = self.data_parallel_size # DP address, used in multi-node case for torch distributed group # and ZMQ sockets. if self.data_parallel_address is None: if self.data_parallel_backend == "ray": host_ip = get_ip() logger.info( "Using host IP %s as ray-based data parallel address", host_ip ) data_parallel_address = host_ip else: assert self.data_parallel_backend == "mp", ( "data_parallel_backend can only be ray or mp, got %s", self.data_parallel_backend, ) data_parallel_address = ( self.master_addr or ParallelConfig.data_parallel_master_ip ) else: data_parallel_address = self.data_parallel_address # This port is only used when there are remote data parallel engines, # otherwise the local IPC transport is used. data_parallel_rpc_port = ( self.data_parallel_rpc_port if (self.data_parallel_rpc_port is not None) else ParallelConfig.data_parallel_rpc_port ) if self.tokens_only and not model_config.skip_tokenizer_init: model_config.skip_tokenizer_init = True logger.info("Skipping tokenizer initialization for tokens-only mode.") parallel_config = ParallelConfig( pipeline_parallel_size=self.pipeline_parallel_size, tensor_parallel_size=self.tensor_parallel_size, prefill_context_parallel_size=self.prefill_context_parallel_size, data_parallel_size=self.data_parallel_size, data_parallel_rank=self.data_parallel_rank or 0, data_parallel_external_lb=data_parallel_external_lb, data_parallel_size_local=data_parallel_size_local, master_addr=self.master_addr, master_port=self.master_port, nnodes=self.nnodes, node_rank=self.node_rank, data_parallel_master_ip=data_parallel_address, data_parallel_rpc_port=data_parallel_rpc_port, data_parallel_backend=self.data_parallel_backend, data_parallel_hybrid_lb=self.data_parallel_hybrid_lb, enable_expert_parallel=self.enable_expert_parallel, all2all_backend=self.all2all_backend, enable_dbo=self.enable_dbo, ubatch_size=self.ubatch_size, dbo_decode_token_threshold=self.dbo_decode_token_threshold, dbo_prefill_token_threshold=self.dbo_prefill_token_threshold, disable_nccl_for_dp_synchronization=self.disable_nccl_for_dp_synchronization, enable_eplb=self.enable_eplb, eplb_config=self.eplb_config, expert_placement_strategy=self.expert_placement_strategy, max_parallel_loading_workers=self.max_parallel_loading_workers, disable_custom_all_reduce=self.disable_custom_all_reduce, ray_workers_use_nsight=self.ray_workers_use_nsight, ray_runtime_env=ray_runtime_env, placement_group=placement_group, distributed_executor_backend=self.distributed_executor_backend, worker_cls=self.worker_cls, worker_extension_cls=self.worker_extension_cls, decode_context_parallel_size=self.decode_context_parallel_size, dcp_kv_cache_interleave_size=self.dcp_kv_cache_interleave_size, cp_kv_cache_interleave_size=self.cp_kv_cache_interleave_size, _api_process_count=self._api_process_count, _api_process_rank=self._api_process_rank, ) speculative_config = self.create_speculative_config( target_model_config=model_config, target_parallel_config=parallel_config, ) scheduler_config = SchedulerConfig( runner_type=model_config.runner_type, max_num_batched_tokens=self.max_num_batched_tokens, max_num_seqs=self.max_num_seqs, max_model_len=model_config.max_model_len, enable_chunked_prefill=self.enable_chunked_prefill, disable_chunked_mm_input=self.disable_chunked_mm_input, is_multimodal_model=model_config.is_multimodal_model, is_encoder_decoder=model_config.is_encoder_decoder, policy=self.scheduling_policy, scheduler_cls=self.scheduler_cls, max_num_partial_prefills=self.max_num_partial_prefills, max_long_partial_prefills=self.max_long_partial_prefills, long_prefill_token_threshold=self.long_prefill_token_threshold, disable_hybrid_kv_cache_manager=self.disable_hybrid_kv_cache_manager, async_scheduling=self.async_scheduling, stream_interval=self.stream_interval, ) if not model_config.is_multimodal_model and self.default_mm_loras: raise ValueError( "Default modality-specific LoRA(s) were provided for a " "non multimodal model" ) lora_config = ( LoRAConfig( max_lora_rank=self.max_lora_rank, max_loras=self.max_loras, default_mm_loras=self.default_mm_loras, fully_sharded_loras=self.fully_sharded_loras, lora_dtype=self.lora_dtype, max_cpu_loras=self.max_cpu_loras if self.max_cpu_loras and self.max_cpu_loras > 0 else None, ) if self.enable_lora else None ) if ( lora_config is not None and speculative_config is not None and scheduler_config.max_num_batched_tokens < ( scheduler_config.max_num_seqs * (speculative_config.num_speculative_tokens + 1) ) ): raise ValueError( "Consider increasing max_num_batched_tokens or " "decreasing num_speculative_tokens" ) # bitsandbytes pre-quantized model need a specific model loader if model_config.quantization == "bitsandbytes": self.quantization = self.load_format = "bitsandbytes" # Attention config overrides attention_config = copy.deepcopy(self.attention_config) if self.attention_backend is not None: if attention_config.backend is not None: raise ValueError( "attention_backend and attention_config.backend " "are mutually exclusive" ) # Convert string to enum if needed (CLI parsing returns a string) if isinstance(self.attention_backend, str): attention_config.backend = AttentionBackendEnum[ self.attention_backend.upper() ] else: attention_config.backend = self.attention_backend load_config = self.create_load_config() # Pass reasoning_parser into StructuredOutputsConfig if self.reasoning_parser: self.structured_outputs_config.reasoning_parser = self.reasoning_parser if self.reasoning_parser_plugin: self.structured_outputs_config.reasoning_parser_plugin = ( self.reasoning_parser_plugin ) observability_config = ObservabilityConfig( show_hidden_metrics_for_version=self.show_hidden_metrics_for_version, otlp_traces_endpoint=self.otlp_traces_endpoint, collect_detailed_traces=self.collect_detailed_traces, kv_cache_metrics=self.kv_cache_metrics, kv_cache_metrics_sample=self.kv_cache_metrics_sample, cudagraph_metrics=self.cudagraph_metrics, enable_layerwise_nvtx_tracing=self.enable_layerwise_nvtx_tracing, enable_mfu_metrics=self.enable_mfu_metrics, ) # Compilation config overrides compilation_config = copy.deepcopy(self.compilation_config) if self.cudagraph_capture_sizes is not None: if compilation_config.cudagraph_capture_sizes is not None: raise ValueError( "cudagraph_capture_sizes and compilation_config." "cudagraph_capture_sizes are mutually exclusive" ) compilation_config.cudagraph_capture_sizes = self.cudagraph_capture_sizes if self.max_cudagraph_capture_size is not None: if compilation_config.max_cudagraph_capture_size is not None: raise ValueError( "max_cudagraph_capture_size and compilation_config." "max_cudagraph_capture_size are mutually exclusive" ) compilation_config.max_cudagraph_capture_size = ( self.max_cudagraph_capture_size ) config = VllmConfig( model_config=model_config, cache_config=cache_config, parallel_config=parallel_config, scheduler_config=scheduler_config, device_config=device_config, load_config=load_config, attention_config=attention_config, lora_config=lora_config, speculative_config=speculative_config, structured_outputs_config=self.structured_outputs_config, observability_config=observability_config, compilation_config=compilation_config, kv_transfer_config=self.kv_transfer_config, kv_events_config=self.kv_events_config, ec_transfer_config=self.ec_transfer_config, profiler_config=self.profiler_config, additional_config=self.additional_config, optimization_level=self.optimization_level, ) return config def _check_feature_supported(self, model_config: ModelConfig): """Raise an error if the feature is not supported.""" if self.logits_processor_pattern != EngineArgs.logits_processor_pattern: _raise_unsupported_error(feature_name="--logits-processor-pattern") # No Concurrent Partial Prefills so far. if ( self.max_num_partial_prefills != SchedulerConfig.max_num_partial_prefills or self.max_long_partial_prefills != SchedulerConfig.max_long_partial_prefills ): _raise_unsupported_error(feature_name="Concurrent Partial Prefill") # N-gram, Medusa, and Eagle are supported for speculative decoding. if self.speculative_config is not None: # speculative_config could still be a dict at this point if isinstance(self.speculative_config, dict): method = self.speculative_config.get("method", None) else: method = self.speculative_config.method if method == "draft_model": raise NotImplementedError( "Draft model speculative decoding is not supported yet. " "Please consider using other speculative decoding methods " "such as ngram, medusa, eagle, or mtp." ) if self.pipeline_parallel_size > 1: supports_pp = getattr( self.distributed_executor_backend, "supports_pp", False ) if not supports_pp and self.distributed_executor_backend not in ( ParallelConfig.distributed_executor_backend, "ray", "mp", "external_launcher", ): name = ( "Pipeline Parallelism without Ray distributed " "executor or multiprocessing executor or external " "launcher" ) _raise_unsupported_error(feature_name=name) @classmethod def get_batch_defaults( cls, world_size: int, ) -> tuple[dict[UsageContext | None, int], dict[UsageContext | None, int]]: from vllm.usage.usage_lib import UsageContext default_max_num_batched_tokens: dict[UsageContext | None, int] default_max_num_seqs: dict[UsageContext | None, int] # When no user override, set the default values based on the usage # context. # Use different default values for different hardware. # Try to query the device name on the current platform. If it fails, # it may be because the platform that imports vLLM is not the same # as the platform that vLLM is running on (e.g. the case of scaling # vLLM with Ray) and has no GPUs. In this case we use the default # values for non-H100/H200 GPUs. try: device_memory = current_platform.get_device_total_memory() device_name = current_platform.get_device_name().lower() except Exception: # This is only used to set default_max_num_batched_tokens device_memory = 0 device_name = "" # NOTE(Kuntai): Setting large `max_num_batched_tokens` for A100 reduces # throughput, see PR #17885 for more details. # So here we do an extra device name check to prevent such regression. if device_memory >= 70 * GiB_bytes and "a100" not in device_name: # For GPUs like H100 and MI300x, use larger default values. default_max_num_batched_tokens = { UsageContext.LLM_CLASS: 16384, UsageContext.OPENAI_API_SERVER: 8192, } default_max_num_seqs = { UsageContext.LLM_CLASS: 1024, UsageContext.OPENAI_API_SERVER: 1024, } else: # TODO(woosuk): Tune the default values for other hardware. default_max_num_batched_tokens = { UsageContext.LLM_CLASS: 8192, UsageContext.OPENAI_API_SERVER: 2048, } default_max_num_seqs = { UsageContext.LLM_CLASS: 256, UsageContext.OPENAI_API_SERVER: 256, } # tpu specific default values. if current_platform.is_tpu(): chip_name = current_platform.get_device_name() if chip_name == "V6E": default_max_num_batched_tokens = { UsageContext.LLM_CLASS: 2048, UsageContext.OPENAI_API_SERVER: 1024, } elif chip_name == "V5E": default_max_num_batched_tokens = { UsageContext.LLM_CLASS: 1024, UsageContext.OPENAI_API_SERVER: 512, } elif chip_name == "V5P": default_max_num_batched_tokens = { UsageContext.LLM_CLASS: 512, UsageContext.OPENAI_API_SERVER: 256, } # cpu specific default values. if current_platform.is_cpu(): default_max_num_batched_tokens = { UsageContext.LLM_CLASS: 4096 * world_size, UsageContext.OPENAI_API_SERVER: 2048 * world_size, } default_max_num_seqs = { UsageContext.LLM_CLASS: 256 * world_size, UsageContext.OPENAI_API_SERVER: 128 * world_size, } return default_max_num_batched_tokens, default_max_num_seqs def _set_default_chunked_prefill_and_prefix_caching_args( self, model_config: ModelConfig ) -> None: default_chunked_prefill = model_config.is_chunked_prefill_supported default_prefix_caching = model_config.is_prefix_caching_supported if self.enable_chunked_prefill is None: self.enable_chunked_prefill = default_chunked_prefill logger.debug( "%s chunked prefill by default", "Enabling" if default_chunked_prefill else "Disabling", ) elif ( model_config.runner_type == "generate" and not self.enable_chunked_prefill and default_chunked_prefill ): logger.warning_once( "This model does not officially support disabling chunked prefill. " "Disabling this manually may cause the engine to crash " "or produce incorrect outputs.", scope="local", ) elif ( model_config.runner_type == "pooling" and self.enable_chunked_prefill and not default_chunked_prefill ): logger.warning_once( "This model does not officially support chunked prefill. " "Enabling this manually may cause the engine to crash " "or produce incorrect outputs.", scope="local", ) if self.enable_prefix_caching is None: self.enable_prefix_caching = default_prefix_caching logger.debug( "%s prefix caching by default", "Enabling" if default_prefix_caching else "Disabling", ) elif ( model_config.runner_type == "pooling" and self.enable_prefix_caching and not default_prefix_caching ): logger.warning_once( "This model does not officially support prefix caching. " "Enabling this manually may cause the engine to crash " "or produce incorrect outputs.", scope="local", ) # Disable chunked prefill and prefix caching for: # POWER (ppc64le)/s390x/RISCV CPUs in V1 if current_platform.is_cpu() and current_platform.get_cpu_architecture() in ( CpuArchEnum.POWERPC, CpuArchEnum.S390X, CpuArchEnum.RISCV, ): logger.info( "Chunked prefill is not supported for ARM and POWER, " "S390X and RISC-V CPUs; " "disabling it for V1 backend." ) self.enable_chunked_prefill = False logger.info( "Prefix caching is not supported for ARM and POWER, " "S390X and RISC-V CPUs; " "disabling it for V1 backend." ) self.enable_prefix_caching = False def _set_default_max_num_seqs_and_batched_tokens_args( self, usage_context: UsageContext | None, model_config: ModelConfig, ): world_size = self.pipeline_parallel_size * self.tensor_parallel_size ( default_max_num_batched_tokens, default_max_num_seqs, ) = self.get_batch_defaults(world_size) orig_max_num_batched_tokens = self.max_num_batched_tokens orig_max_num_seqs = self.max_num_seqs if self.max_num_batched_tokens is None: self.max_num_batched_tokens = default_max_num_batched_tokens.get( usage_context, SchedulerConfig.DEFAULT_MAX_NUM_BATCHED_TOKENS, ) if self.max_num_seqs is None: self.max_num_seqs = default_max_num_seqs.get( usage_context, SchedulerConfig.DEFAULT_MAX_NUM_SEQS, ) if orig_max_num_batched_tokens is None: if not self.enable_chunked_prefill: # If max_model_len is too short, use the default for higher throughput. self.max_num_batched_tokens = max( model_config.max_model_len, self.max_num_batched_tokens, ) # When using default settings, # Ensure max_num_batched_tokens does not exceed model limit. # Some models (e.g., Whisper) have embeddings tied to max length. self.max_num_batched_tokens = min( self.max_num_seqs * model_config.max_model_len, self.max_num_batched_tokens, ) logger.debug( "Defaulting max_num_batched_tokens to %d for %s usage context.", self.max_num_batched_tokens, usage_context.value if usage_context else None, ) if orig_max_num_seqs is None: assert self.max_num_batched_tokens is not None # For type checking self.max_num_seqs = min(self.max_num_seqs, self.max_num_batched_tokens) logger.debug( "Defaulting max_num_seqs to %d for %s usage context.", self.max_num_seqs, usage_context.value if usage_context else None, ) _api_process_count class-attribute instance-attribute ¶ _api_process_count: int = _api_process_count _api_process_rank class-attribute instance-attribute ¶ _api_process_rank: int = _api_process_rank additional_config class-attribute instance-attribute ¶ additional_config: dict[str, Any] = get_field( VllmConfig, "additional_config" ) aggregate_engine_logging class-attribute instance-attribute ¶ aggregate_engine_logging: bool = False all2all_backend class-attribute instance-attribute ¶ all2all_backend: str = all2all_backend allowed_local_media_path class-attribute instance-attribute ¶ allowed_local_media_path: str = allowed_local_media_path allowed_media_domains class-attribute instance-attribute ¶ allowed_media_domains: list[str] | None = ( allowed_media_domains ) async_scheduling class-attribute instance-attribute ¶ async_scheduling: bool | None = async_scheduling attention_backend class-attribute instance-attribute ¶ attention_backend: AttentionBackendEnum | None = backend attention_config class-attribute instance-attribute ¶ attention_config: AttentionConfig = get_field( VllmConfig, "attention_config" ) block_size class-attribute instance-attribute ¶ block_size: BlockSize | None = block_size calculate_kv_scales class-attribute instance-attribute ¶ calculate_kv_scales: bool = calculate_kv_scales code_revision class-attribute instance-attribute ¶ code_revision: str | None = code_revision collect_detailed_traces class-attribute instance-attribute ¶ collect_detailed_traces: ( list[DetailedTraceModules] | None ) = collect_detailed_traces compilation_config class-attribute instance-attribute ¶ compilation_config: CompilationConfig = get_field( VllmConfig, "compilation_config" ) config_format class-attribute instance-attribute ¶ config_format: str = config_format convert class-attribute instance-attribute ¶ convert: ConvertOption = convert cp_kv_cache_interleave_size class-attribute instance-attribute ¶ cp_kv_cache_interleave_size: int = ( cp_kv_cache_interleave_size ) cpu_offload_gb class-attribute instance-attribute ¶ cpu_offload_gb: float = cpu_offload_gb cudagraph_capture_sizes class-attribute instance-attribute ¶ cudagraph_capture_sizes: list[int] | None = ( cudagraph_capture_sizes ) cudagraph_metrics class-attribute instance-attribute ¶ cudagraph_metrics: bool = cudagraph_metrics data_parallel_address class-attribute instance-attribute ¶ data_parallel_address: str | None = None data_parallel_backend class-attribute instance-attribute ¶ data_parallel_backend: str = data_parallel_backend data_parallel_external_lb class-attribute instance-attribute ¶ data_parallel_external_lb: bool = False data_parallel_hybrid_lb class-attribute instance-attribute ¶ data_parallel_hybrid_lb: bool = False data_parallel_rank class-attribute instance-attribute ¶ data_parallel_rank: int | None = None data_parallel_rpc_port class-attribute instance-attribute ¶ data_parallel_rpc_port: int | None = None data_parallel_size class-attribute instance-attribute ¶ data_parallel_size: int = data_parallel_size data_parallel_size_local class-attribute instance-attribute ¶ data_parallel_size_local: int | None = None data_parallel_start_rank class-attribute instance-attribute ¶ data_parallel_start_rank: int | None = None dbo_decode_token_threshold class-attribute instance-attribute ¶ dbo_decode_token_threshold: int = dbo_decode_token_threshold dbo_prefill_token_threshold class-attribute instance-attribute ¶ dbo_prefill_token_threshold: int = ( dbo_prefill_token_threshold ) dcp_kv_cache_interleave_size class-attribute instance-attribute ¶ dcp_kv_cache_interleave_size: int = ( dcp_kv_cache_interleave_size ) decode_context_parallel_size class-attribute instance-attribute ¶ decode_context_parallel_size: int = ( decode_context_parallel_size ) default_mm_loras class-attribute instance-attribute ¶ default_mm_loras: dict[str, str] | None = default_mm_loras disable_cascade_attn class-attribute instance-attribute ¶ disable_cascade_attn: bool = disable_cascade_attn disable_chunked_mm_input class-attribute instance-attribute ¶ disable_chunked_mm_input: bool = disable_chunked_mm_input disable_custom_all_reduce class-attribute instance-attribute ¶ disable_custom_all_reduce: bool = disable_custom_all_reduce disable_hybrid_kv_cache_manager class-attribute instance-attribute ¶ disable_hybrid_kv_cache_manager: bool | None = ( disable_hybrid_kv_cache_manager ) disable_log_stats class-attribute instance-attribute ¶ disable_log_stats: bool = False disable_nccl_for_dp_synchronization class-attribute instance-attribute ¶ disable_nccl_for_dp_synchronization: bool = ( disable_nccl_for_dp_synchronization ) disable_sliding_window class-attribute instance-attribute ¶ disable_sliding_window: bool = disable_sliding_window distributed_executor_backend class-attribute instance-attribute ¶ distributed_executor_backend: ( str | DistributedExecutorBackend | type[Executor] | None ) = distributed_executor_backend download_dir class-attribute instance-attribute ¶ download_dir: str | None = download_dir dtype class-attribute instance-attribute ¶ dtype: ModelDType = dtype ec_transfer_config class-attribute instance-attribute ¶ ec_transfer_config: ECTransferConfig | None = None enable_chunked_prefill class-attribute instance-attribute ¶ enable_chunked_prefill: bool | None = None enable_dbo class-attribute instance-attribute ¶ enable_dbo: bool = enable_dbo enable_eplb class-attribute instance-attribute ¶ enable_eplb: bool = enable_eplb enable_expert_parallel class-attribute instance-attribute ¶ enable_expert_parallel: bool = enable_expert_parallel enable_layerwise_nvtx_tracing class-attribute instance-attribute ¶ enable_layerwise_nvtx_tracing: bool = ( enable_layerwise_nvtx_tracing ) enable_lora class-attribute instance-attribute ¶ enable_lora: bool = False enable_mfu_metrics class-attribute instance-attribute ¶ enable_mfu_metrics: bool = enable_mfu_metrics enable_mm_embeds class-attribute instance-attribute ¶ enable_mm_embeds: bool = enable_mm_embeds enable_prefix_caching class-attribute instance-attribute ¶ enable_prefix_caching: bool | None = None enable_prompt_embeds class-attribute instance-attribute ¶ enable_prompt_embeds: bool = enable_prompt_embeds enable_sleep_mode class-attribute instance-attribute ¶ enable_sleep_mode: bool = enable_sleep_mode enforce_eager class-attribute instance-attribute ¶ enforce_eager: bool = enforce_eager eplb_config class-attribute instance-attribute ¶ eplb_config: EPLBConfig = get_field( ParallelConfig, "eplb_config" ) expert_placement_strategy class-attribute instance-attribute ¶ expert_placement_strategy: ExpertPlacementStrategy = ( expert_placement_strategy ) fully_sharded_loras class-attribute instance-attribute ¶ fully_sharded_loras: bool = fully_sharded_loras generation_config class-attribute instance-attribute ¶ generation_config: str = generation_config gpu_memory_utilization class-attribute instance-attribute ¶ gpu_memory_utilization: float = gpu_memory_utilization hf_config_path class-attribute instance-attribute ¶ hf_config_path: str | None = hf_config_path hf_overrides class-attribute instance-attribute ¶ hf_overrides: HfOverrides = get_field( ModelConfig, "hf_overrides" ) hf_token class-attribute instance-attribute ¶ hf_token: bool | str | None = hf_token ignore_patterns class-attribute instance-attribute ¶ ignore_patterns: str | list[str] = get_field( LoadConfig, "ignore_patterns" ) interleave_mm_strings class-attribute instance-attribute ¶ interleave_mm_strings: bool = interleave_mm_strings io_processor_plugin class-attribute instance-attribute ¶ io_processor_plugin: str | None = None kv_cache_dtype class-attribute instance-attribute ¶ kv_cache_dtype: CacheDType = cache_dtype kv_cache_memory_bytes class-attribute instance-attribute ¶ kv_cache_memory_bytes: int | None = kv_cache_memory_bytes kv_cache_metrics class-attribute instance-attribute ¶ kv_cache_metrics: bool = kv_cache_metrics kv_cache_metrics_sample class-attribute instance-attribute ¶ kv_cache_metrics_sample: float = get_field( ObservabilityConfig, "kv_cache_metrics_sample" ) kv_events_config class-attribute instance-attribute ¶ kv_events_config: KVEventsConfig | None = None kv_offloading_backend class-attribute instance-attribute ¶ kv_offloading_backend: KVOffloadingBackend | None = ( kv_offloading_backend ) kv_offloading_size class-attribute instance-attribute ¶ kv_offloading_size: float | None = kv_offloading_size kv_sharing_fast_prefill class-attribute instance-attribute ¶ kv_sharing_fast_prefill: bool = kv_sharing_fast_prefill kv_transfer_config class-attribute instance-attribute ¶ kv_transfer_config: KVTransferConfig | None = None limit_mm_per_prompt class-attribute instance-attribute ¶ limit_mm_per_prompt: dict[str, int | dict[str, int]] = ( get_field(MultiModalConfig, "limit_per_prompt") ) load_format class-attribute instance-attribute ¶ load_format: str | LoadFormats = load_format logits_processor_pattern class-attribute instance-attribute ¶ logits_processor_pattern: str | None = ( logits_processor_pattern ) logits_processors class-attribute instance-attribute ¶ logits_processors: ( list[str | type[LogitsProcessor]] | None ) = logits_processors Custom logitproc types logprobs_mode class-attribute instance-attribute ¶ logprobs_mode: LogprobsMode = logprobs_mode long_prefill_token_threshold class-attribute instance-attribute ¶ long_prefill_token_threshold: int = ( long_prefill_token_threshold ) lora_dtype class-attribute instance-attribute ¶ lora_dtype: str | dtype | None = lora_dtype mamba_block_size class-attribute instance-attribute ¶ mamba_block_size: int | None = get_field( CacheConfig, "mamba_block_size" ) mamba_cache_dtype class-attribute instance-attribute ¶ mamba_cache_dtype: MambaDType = mamba_cache_dtype mamba_ssm_cache_dtype class-attribute instance-attribute ¶ mamba_ssm_cache_dtype: MambaDType = mamba_ssm_cache_dtype master_addr class-attribute instance-attribute ¶ master_addr: str = master_addr master_port class-attribute instance-attribute ¶ master_port: int = master_port max_cpu_loras class-attribute instance-attribute ¶ max_cpu_loras: int | None = max_cpu_loras max_cudagraph_capture_size class-attribute instance-attribute ¶ max_cudagraph_capture_size: int | None = get_field( CompilationConfig, "max_cudagraph_capture_size" ) max_logprobs class-attribute instance-attribute ¶ max_logprobs: int = max_logprobs max_long_partial_prefills class-attribute instance-attribute ¶ max_long_partial_prefills: int = max_long_partial_prefills max_lora_rank class-attribute instance-attribute ¶ max_lora_rank: int = max_lora_rank max_loras class-attribute instance-attribute ¶ max_loras: int = max_loras max_model_len class-attribute instance-attribute ¶ max_model_len: int | None = max_model_len max_num_batched_tokens class-attribute instance-attribute ¶ max_num_batched_tokens: int | None = None max_num_partial_prefills class-attribute instance-attribute ¶ max_num_partial_prefills: int = max_num_partial_prefills max_num_seqs class-attribute instance-attribute ¶ max_num_seqs: int | None = None max_parallel_loading_workers class-attribute instance-attribute ¶ max_parallel_loading_workers: int | None = ( max_parallel_loading_workers ) media_io_kwargs class-attribute instance-attribute ¶ media_io_kwargs: dict[str, dict[str, Any]] = get_field( MultiModalConfig, "media_io_kwargs" ) mm_encoder_attn_backend class-attribute instance-attribute ¶ mm_encoder_attn_backend: ( AttentionBackendEnum | str | None ) = mm_encoder_attn_backend mm_encoder_tp_mode class-attribute instance-attribute ¶ mm_encoder_tp_mode: MMEncoderTPMode = mm_encoder_tp_mode mm_processor_cache_gb class-attribute instance-attribute ¶ mm_processor_cache_gb: float = mm_processor_cache_gb mm_processor_cache_type class-attribute instance-attribute ¶ mm_processor_cache_type: MMCacheType | None = ( mm_processor_cache_type ) mm_processor_kwargs class-attribute instance-attribute ¶ mm_processor_kwargs: dict[str, Any] | None = ( mm_processor_kwargs ) mm_shm_cache_max_object_size_mb class-attribute instance-attribute ¶ mm_shm_cache_max_object_size_mb: int = ( mm_shm_cache_max_object_size_mb ) model class-attribute instance-attribute ¶ model: str = model model_impl class-attribute instance-attribute ¶ model_impl: str = model_impl model_loader_extra_config class-attribute instance-attribute ¶ model_loader_extra_config: dict = get_field( LoadConfig, "model_loader_extra_config" ) nnodes class-attribute instance-attribute ¶ nnodes: int = nnodes node_rank class-attribute instance-attribute ¶ node_rank: int = node_rank num_gpu_blocks_override class-attribute instance-attribute ¶ num_gpu_blocks_override: int | None = ( num_gpu_blocks_override ) optimization_level class-attribute instance-attribute ¶ optimization_level: OptimizationLevel = optimization_level otlp_traces_endpoint class-attribute instance-attribute ¶ otlp_traces_endpoint: str | None = otlp_traces_endpoint override_attention_dtype class-attribute instance-attribute ¶ override_attention_dtype: str = override_attention_dtype override_generation_config class-attribute instance-attribute ¶ override_generation_config: dict[str, Any] = get_field( ModelConfig, "override_generation_config" ) pipeline_parallel_size class-attribute instance-attribute ¶ pipeline_parallel_size: int = pipeline_parallel_size pooler_config class-attribute instance-attribute ¶ pooler_config: PoolerConfig | None = pooler_config prefill_context_parallel_size class-attribute instance-attribute ¶ prefill_context_parallel_size: int = ( prefill_context_parallel_size ) prefix_caching_hash_algo class-attribute instance-attribute ¶ prefix_caching_hash_algo: PrefixCachingHashAlgo = ( prefix_caching_hash_algo ) profiler_config class-attribute instance-attribute ¶ profiler_config: ProfilerConfig = get_field( VllmConfig, "profiler_config" ) pt_load_map_location class-attribute instance-attribute ¶ pt_load_map_location: str = pt_load_map_location quantization class-attribute instance-attribute ¶ quantization: QuantizationMethods | None = quantization ray_workers_use_nsight class-attribute instance-attribute ¶ ray_workers_use_nsight: bool = ray_workers_use_nsight reasoning_parser class-attribute instance-attribute ¶ reasoning_parser: str = reasoning_parser reasoning_parser_plugin class-attribute instance-attribute ¶ reasoning_parser_plugin: str | None = None revision class-attribute instance-attribute ¶ revision: str | None = revision runner class-attribute instance-attribute ¶ runner: RunnerOption = runner safetensors_load_strategy class-attribute instance-attribute ¶ safetensors_load_strategy: str = safetensors_load_strategy scheduler_cls class-attribute instance-attribute ¶ scheduler_cls: str | type[object] | None = scheduler_cls scheduling_policy class-attribute instance-attribute ¶ scheduling_policy: SchedulerPolicy = policy seed class-attribute instance-attribute ¶ seed: int = seed served_model_name class-attribute instance-attribute ¶ served_model_name: str | list[str] | None = ( served_model_name ) show_hidden_metrics_for_version class-attribute instance-attribute ¶ show_hidden_metrics_for_version: str | None = ( show_hidden_metrics_for_version ) skip_mm_profiling class-attribute instance-attribute ¶ skip_mm_profiling: bool = skip_mm_profiling skip_tokenizer_init class-attribute instance-attribute ¶ skip_tokenizer_init: bool = skip_tokenizer_init speculative_config class-attribute instance-attribute ¶ speculative_config: dict[str, Any] | None = None stream_interval class-attribute instance-attribute ¶ stream_interval: int = stream_interval structured_outputs_config class-attribute instance-attribute ¶ structured_outputs_config: StructuredOutputsConfig = ( get_field(VllmConfig, "structured_outputs_config") ) swap_space class-attribute instance-attribute ¶ swap_space: float = swap_space tensor_parallel_size class-attribute instance-attribute ¶ tensor_parallel_size: int = tensor_parallel_size tokenizer class-attribute instance-attribute ¶ tokenizer: str | None = tokenizer tokenizer_mode class-attribute instance-attribute ¶ tokenizer_mode: TokenizerMode | str = tokenizer_mode tokenizer_revision class-attribute instance-attribute ¶ tokenizer_revision: str | None = tokenizer_revision tokens_only class-attribute instance-attribute ¶ tokens_only: bool = False trust_remote_code class-attribute instance-attribute ¶ trust_remote_code: bool = trust_remote_code ubatch_size class-attribute instance-attribute ¶ ubatch_size: int = ubatch_size use_tqdm_on_load class-attribute instance-attribute ¶ use_tqdm_on_load: bool = use_tqdm_on_load video_pruning_rate class-attribute instance-attribute ¶ video_pruning_rate: float = video_pruning_rate worker_cls class-attribute instance-attribute ¶ worker_cls: str = worker_cls worker_extension_cls class-attribute instance-attribute ¶ worker_extension_cls: str = worker_extension_cls __init__ ¶ __init__( model: str = model, served_model_name: str | list[str] | None = served_model_name, tokenizer: str | None = tokenizer, hf_config_path: str | None = hf_config_path, runner: RunnerOption = runner, convert: ConvertOption = convert, skip_tokenizer_init: bool = skip_tokenizer_init, enable_prompt_embeds: bool = enable_prompt_embeds, tokenizer_mode: TokenizerMode | str = tokenizer_mode, trust_remote_code: bool = trust_remote_code, allowed_local_media_path: str = allowed_local_media_path, allowed_media_domains: list[str] | None = allowed_media_domains, download_dir: str | None = download_dir, safetensors_load_strategy: str = safetensors_load_strategy, load_format: str | LoadFormats = load_format, config_format: str = config_format, dtype: ModelDType = dtype, kv_cache_dtype: CacheDType = cache_dtype, seed: int = seed, max_model_len: int | None = max_model_len, cudagraph_capture_sizes: list[int] | None = cudagraph_capture_sizes, max_cudagraph_capture_size: int | None = get_field( CompilationConfig, "max_cudagraph_capture_size" ), distributed_executor_backend: str | DistributedExecutorBackend | type[Executor] | None = distributed_executor_backend, pipeline_parallel_size: int = pipeline_parallel_size, master_addr: str = master_addr, master_port: int = master_port, nnodes: int = nnodes, node_rank: int = node_rank, tensor_parallel_size: int = tensor_parallel_size, prefill_context_parallel_size: int = prefill_context_parallel_size, decode_context_parallel_size: int = decode_context_parallel_size, dcp_kv_cache_interleave_size: int = dcp_kv_cache_interleave_size, cp_kv_cache_interleave_size: int = cp_kv_cache_interleave_size, data_parallel_size: int = data_parallel_size, data_parallel_rank: int | None = None, data_parallel_start_rank: int | None = None, data_parallel_size_local: int | None = None, data_parallel_address: str | None = None, data_parallel_rpc_port: int | None = None, data_parallel_hybrid_lb: bool = False, data_parallel_external_lb: bool = False, data_parallel_backend: str = data_parallel_backend, enable_expert_parallel: bool = enable_expert_parallel, all2all_backend: str = all2all_backend, enable_dbo: bool = enable_dbo, ubatch_size: int = ubatch_size, dbo_decode_token_threshold: int = dbo_decode_token_threshold, dbo_prefill_token_threshold: int = dbo_prefill_token_threshold, disable_nccl_for_dp_synchronization: bool = disable_nccl_for_dp_synchronization, eplb_config: EPLBConfig = get_field( ParallelConfig, "eplb_config" ), enable_eplb: bool = enable_eplb, expert_placement_strategy: ExpertPlacementStrategy = expert_placement_strategy, _api_process_count: int = _api_process_count, _api_process_rank: int = _api_process_rank, max_parallel_loading_workers: int | None = max_parallel_loading_workers, block_size: BlockSize | None = block_size, enable_prefix_caching: bool | None = None, prefix_caching_hash_algo: PrefixCachingHashAlgo = prefix_caching_hash_algo, disable_sliding_window: bool = disable_sliding_window, disable_cascade_attn: bool = disable_cascade_attn, swap_space: float = swap_space, cpu_offload_gb: float = cpu_offload_gb, gpu_memory_utilization: float = gpu_memory_utilization, kv_cache_memory_bytes: int | None = kv_cache_memory_bytes, max_num_batched_tokens: int | None = None, max_num_partial_prefills: int = max_num_partial_prefills, max_long_partial_prefills: int = max_long_partial_prefills, long_prefill_token_threshold: int = long_prefill_token_threshold, max_num_seqs: int | None = None, max_logprobs: int = max_logprobs, logprobs_mode: LogprobsMode = logprobs_mode, disable_log_stats: bool = False, aggregate_engine_logging: bool = False, revision: str | None = revision, code_revision: str | None = code_revision, hf_token: bool | str | None = hf_token, hf_overrides: HfOverrides = get_field( ModelConfig, "hf_overrides" ), tokenizer_revision: str | None = tokenizer_revision, quantization: QuantizationMethods | None = quantization, enforce_eager: bool = enforce_eager, disable_custom_all_reduce: bool = disable_custom_all_reduce, limit_mm_per_prompt: dict[ str, int | dict[str, int] ] = get_field(MultiModalConfig, "limit_per_prompt"), enable_mm_embeds: bool = enable_mm_embeds, interleave_mm_strings: bool = interleave_mm_strings, media_io_kwargs: dict[str, dict[str, Any]] = get_field( MultiModalConfig, "media_io_kwargs" ), mm_processor_kwargs: dict[str, Any] | None = mm_processor_kwargs, mm_processor_cache_gb: float = mm_processor_cache_gb, mm_processor_cache_type: MMCacheType | None = mm_processor_cache_type, mm_shm_cache_max_object_size_mb: int = mm_shm_cache_max_object_size_mb, mm_encoder_tp_mode: MMEncoderTPMode = mm_encoder_tp_mode, mm_encoder_attn_backend: AttentionBackendEnum | str | None = mm_encoder_attn_backend, io_processor_plugin: str | None = None, skip_mm_profiling: bool = skip_mm_profiling, video_pruning_rate: float = video_pruning_rate, enable_lora: bool = False, max_loras: int = max_loras, max_lora_rank: int = max_lora_rank, default_mm_loras: dict[str, str] | None = default_mm_loras, fully_sharded_loras: bool = fully_sharded_loras, max_cpu_loras: int | None = max_cpu_loras, lora_dtype: str | dtype | None = lora_dtype, ray_workers_use_nsight: bool = ray_workers_use_nsight, num_gpu_blocks_override: int | None = num_gpu_blocks_override, model_loader_extra_config: dict = get_field( LoadConfig, "model_loader_extra_config" ), ignore_patterns: str | list[str] = get_field( LoadConfig, "ignore_patterns" ), enable_chunked_prefill: bool | None = None, disable_chunked_mm_input: bool = disable_chunked_mm_input, disable_hybrid_kv_cache_manager: bool | None = disable_hybrid_kv_cache_manager, structured_outputs_config: StructuredOutputsConfig = get_field( VllmConfig, "structured_outputs_config" ), reasoning_parser: str = reasoning_parser, reasoning_parser_plugin: str | None = None, logits_processor_pattern: str | None = logits_processor_pattern, speculative_config: dict[str, Any] | None = None, show_hidden_metrics_for_version: str | None = show_hidden_metrics_for_version, otlp_traces_endpoint: str | None = otlp_traces_endpoint, collect_detailed_traces: list[DetailedTraceModules] | None = collect_detailed_traces, kv_cache_metrics: bool = kv_cache_metrics, kv_cache_metrics_sample: float = get_field( ObservabilityConfig, "kv_cache_metrics_sample" ), cudagraph_metrics: bool = cudagraph_metrics, enable_layerwise_nvtx_tracing: bool = enable_layerwise_nvtx_tracing, enable_mfu_metrics: bool = enable_mfu_metrics, scheduling_policy: SchedulerPolicy = policy, scheduler_cls: str | type[object] | None = scheduler_cls, pooler_config: PoolerConfig | None = pooler_config, compilation_config: CompilationConfig = get_field( VllmConfig, "compilation_config" ), attention_config: AttentionConfig = get_field( VllmConfig, "attention_config" ), worker_cls: str = worker_cls, worker_extension_cls: str = worker_extension_cls, profiler_config: ProfilerConfig = get_field( VllmConfig, "profiler_config" ), kv_transfer_config: KVTransferConfig | None = None, kv_events_config: KVEventsConfig | None = None, ec_transfer_config: ECTransferConfig | None = None, generation_config: str = generation_config, enable_sleep_mode: bool = enable_sleep_mode, override_generation_config: dict[str, Any] = get_field( ModelConfig, "override_generation_config" ), model_impl: str = model_impl, override_attention_dtype: str = override_attention_dtype, attention_backend: AttentionBackendEnum | None = backend, calculate_kv_scales: bool = calculate_kv_scales, mamba_cache_dtype: MambaDType = mamba_cache_dtype, mamba_ssm_cache_dtype: MambaDType = mamba_ssm_cache_dtype, mamba_block_size: int | None = get_field( CacheConfig, "mamba_block_size" ), additional_config: dict[str, Any] = get_field( VllmConfig, "additional_config" ), use_tqdm_on_load: bool = use_tqdm_on_load, pt_load_map_location: str = pt_load_map_location, logits_processors: list[str | type[LogitsProcessor]] | None = logits_processors, async_scheduling: bool | None = async_scheduling, stream_interval: int = stream_interval, kv_sharing_fast_prefill: bool = kv_sharing_fast_prefill, optimization_level: OptimizationLevel = optimization_level, kv_offloading_size: float | None = kv_offloading_size, kv_offloading_backend: KVOffloadingBackend | None = kv_offloading_backend, tokens_only: bool = False, ) -> None __post_init__ ¶ __post_init__() Source code in vllm/engine/arg_utils.py 578 579 580 581 582 583 584 585 586 587 588 589 590 591 592 593 594 595 596 597 598 599 600 601 602 603 604 605 606 607 608 609 610 611def __post_init__(self): # support `EngineArgs(compilation_config={...})` # without having to manually construct a # CompilationConfig object if isinstance(self.compilation_config, dict): self.compilation_config = CompilationConfig(**self.compilation_config) if isinstance(self.attention_config, dict): self.attention_config = AttentionConfig(**self.attention_config) if isinstance(self.eplb_config, dict): self.eplb_config = EPLBConfig(**self.eplb_config) # Setup plugins from vllm.plugins import load_general_plugins load_general_plugins() # when use hf offline,replace model and tokenizer id to local model path if huggingface_hub.constants.HF_HUB_OFFLINE: model_id = self.model self.model = get_model_path(self.model, self.revision) if model_id is not self.model: logger.info( "HF_HUB_OFFLINE is True, replace model_id [%s] to model_path [%s]", model_id, self.model, ) if self.tokenizer is not None: tokenizer_id = self.tokenizer self.tokenizer = get_model_path(self.tokenizer, self.tokenizer_revision) if tokenizer_id is not self.tokenizer: logger.info( "HF_HUB_OFFLINE is True, replace tokenizer_id [%s] " "to tokenizer_path [%s]", tokenizer_id, self.tokenizer, ) _check_feature_supported ¶ _check_feature_supported(model_config: ModelConfig) Raise an error if the feature is not supported. Source code in vllm/engine/arg_utils.py 1739 1740 1741 1742 1743 1744 1745 1746 1747 1748 1749 1750 1751 1752 1753 1754 1755 1756 1757 1758 1759 1760 1761 1762 1763 1764 1765 1766 1767 1768 1769 1770 1771 1772 1773 1774 1775 1776 1777 1778 1779 1780 1781 1782def _check_feature_supported(self, model_config: ModelConfig): """Raise an error if the feature is not supported.""" if self.logits_processor_pattern != EngineArgs.logits_processor_pattern: _raise_unsupported_error(feature_name="--logits-processor-pattern") # No Concurrent Partial Prefills so far. if ( self.max_num_partial_prefills != SchedulerConfig.max_num_partial_prefills or self.max_long_partial_prefills != SchedulerConfig.max_long_partial_prefills ): _raise_unsupported_error(feature_name="Concurrent Partial Prefill") # N-gram, Medusa, and Eagle are supported for speculative decoding. if self.speculative_config is not None: # speculative_config could still be a dict at this point if isinstance(self.speculative_config, dict): method = self.speculative_config.get("method", None) else: method = self.speculative_config.method if method == "draft_model": raise NotImplementedError( "Draft model speculative decoding is not supported yet. " "Please consider using other speculative decoding methods " "such as ngram, medusa, eagle, or mtp." ) if self.pipeline_parallel_size > 1: supports_pp = getattr( self.distributed_executor_backend, "supports_pp", False ) if not supports_pp and self.distributed_executor_backend not in ( ParallelConfig.distributed_executor_backend, "ray", "mp", "external_launcher", ): name = ( "Pipeline Parallelism without Ray distributed " "executor or multiprocessing executor or external " "launcher" ) _raise_unsupported_error(feature_name=name) _set_default_chunked_prefill_and_prefix_caching_args ¶ _set_default_chunked_prefill_and_prefix_caching_args( model_config: ModelConfig, ) -> None Source code in vllm/engine/arg_utils.py 1868 1869 1870 1871 1872 1873 1874 1875 1876 1877 1878 1879 1880 1881 1882 1883 1884 1885 1886 1887 1888 1889 1890 1891 1892 1893 1894 1895 1896 1897 1898 1899 1900 1901 1902 1903 1904 1905 1906 1907 1908 1909 1910 1911 1912 1913 1914 1915 1916 1917 1918 1919 1920 1921 1922 1923 1924 1925 1926 1927 1928 1929 1930 1931 1932 1933 1934 1935 1936 1937 1938 1939 1940 1941def _set_default_chunked_prefill_and_prefix_caching_args( self, model_config: ModelConfig ) -> None: default_chunked_prefill = model_config.is_chunked_prefill_supported default_prefix_caching = model_config.is_prefix_caching_supported if self.enable_chunked_prefill is None: self.enable_chunked_prefill = default_chunked_prefill logger.debug( "%s chunked prefill by default", "Enabling" if default_chunked_prefill else "Disabling", ) elif ( model_config.runner_type == "generate" and not self.enable_chunked_prefill and default_chunked_prefill ): logger.warning_once( "This model does not officially support disabling chunked prefill. " "Disabling this manually may cause the engine to crash " "or produce incorrect outputs.", scope="local", ) elif ( model_config.runner_type == "pooling" and self.enable_chunked_prefill and not default_chunked_prefill ): logger.warning_once( "This model does not officially support chunked prefill. " "Enabling this manually may cause the engine to crash " "or produce incorrect outputs.", scope="local", ) if self.enable_prefix_caching is None: self.enable_prefix_caching = default_prefix_caching logger.debug( "%s prefix caching by default", "Enabling" if default_prefix_caching else "Disabling", ) elif ( model_config.runner_type == "pooling" and self.enable_prefix_caching and not default_prefix_caching ): logger.warning_once( "This model does not officially support prefix caching. " "Enabling this manually may cause the engine to crash " "or produce incorrect outputs.", scope="local", ) # Disable chunked prefill and prefix caching for: # POWER (ppc64le)/s390x/RISCV CPUs in V1 if current_platform.is_cpu() and current_platform.get_cpu_architecture() in ( CpuArchEnum.POWERPC, CpuArchEnum.S390X, CpuArchEnum.RISCV, ): logger.info( "Chunked prefill is not supported for ARM and POWER, " "S390X and RISC-V CPUs; " "disabling it for V1 backend." ) self.enable_chunked_prefill = False logger.info( "Prefix caching is not supported for ARM and POWER, " "S390X and RISC-V CPUs; " "disabling it for V1 backend." ) self.enable_prefix_caching = False _set_default_max_num_seqs_and_batched_tokens_args ¶ _set_default_max_num_seqs_and_batched_tokens_args( usage_context: UsageContext | None, model_config: ModelConfig, ) Source code in vllm/engine/arg_utils.py 1943 1944 1945 1946 1947 1948 1949 1950 1951 1952 1953 1954 1955 1956 1957 1958 1959 1960 1961 1962 1963 1964 1965 1966 1967 1968 1969 1970 1971 1972 1973 1974 1975 1976 1977 1978 1979 1980 1981 1982 1983 1984 1985 1986 1987 1988 1989 1990 1991 1992 1993 1994 1995 1996 1997 1998 1999def _set_default_max_num_seqs_and_batched_tokens_args( self, usage_context: UsageContext | None, model_config: ModelConfig, ): world_size = self.pipeline_parallel_size * self.tensor_parallel_size ( default_max_num_batched_tokens, default_max_num_seqs, ) = self.get_batch_defaults(world_size) orig_max_num_batched_tokens = self.max_num_batched_tokens orig_max_num_seqs = self.max_num_seqs if self.max_num_batched_tokens is None: self.max_num_batched_tokens = default_max_num_batched_tokens.get( usage_context, SchedulerConfig.DEFAULT_MAX_NUM_BATCHED_TOKENS, ) if self.max_num_seqs is None: self.max_num_seqs = default_max_num_seqs.get( usage_context, SchedulerConfig.DEFAULT_MAX_NUM_SEQS, ) if orig_max_num_batched_tokens is None: if not self.enable_chunked_prefill: # If max_model_len is too short, use the default for higher throughput. self.max_num_batched_tokens = max( model_config.max_model_len, self.max_num_batched_tokens, ) # When using default settings, # Ensure max_num_batched_tokens does not exceed model limit. # Some models (e.g., Whisper) have embeddings tied to max length. self.max_num_batched_tokens = min( self.max_num_seqs * model_config.max_model_len, self.max_num_batched_tokens, ) logger.debug( "Defaulting max_num_batched_tokens to %d for %s usage context.", self.max_num_batched_tokens, usage_context.value if usage_context else None, ) if orig_max_num_seqs is None: assert self.max_num_batched_tokens is not None # For type checking self.max_num_seqs = min(self.max_num_seqs, self.max_num_batched_tokens) logger.debug( "Defaulting max_num_seqs to %d for %s usage context.", self.max_num_seqs, usage_context.value if usage_context else None, ) add_cli_args staticmethod ¶ add_cli_args( parser: FlexibleArgumentParser, ) -> FlexibleArgumentParser Shared CLI arguments for vLLM engine. Source code in vllm/engine/arg_utils.py 613 614 615 616 617 618 619 620 621 622 623 624 625 626 627 628 629 630 631 632 633 634 635 636 637 638 639 640 641 642 643 644 645 646 647 648 649 650 651 652 653 654 655 656 657 658 659 660 661 662 663 664 665 666 667 668 669 670 671 672 673 674 675 676 677 678 679 680 681 682 683 684 685 686 687 688 689 690 691 692 693 694 695 696 697 698 699 700 701 702 703 704 705 706 707 708 709 710 711 712 713 714 715 716 717 718 719 720 721 722 723 724 725 726 727 728 729 730 731 732 733 734 735 736 737 738 739 740 741 742 743 744 745 746 747 748 749 750 751 752 753 754 755 756 757 758 759 760 761 762 763 764 765 766 767 768 769 770 771 772 773 774 775 776 777 778 779 780 781 782 783 784 785 786 787 788 789 790 791 792 793 794 795 796 797 798 799 800 801 802 803 804 805 806 807 808 809 810 811 812 813 814 815 816 817 818 819 820 821 822 823 824 825 826 827 828 829 830 831 832 833 834 835 836 837 838 839 840 841 842 843 844 845 846 847 848 849 850 851 852 853 854 855 856 857 858 859 860 861 862 863 864 865 866 867 868 869 870 871 872 873 874 875 876 877 878 879 880 881 882 883 884 885 886 887 888 889 890 891 892 893 894 895 896 897 898 899 900 901 902 903 904 905 906 907 908 909 910 911 912 913 914 915 916 917 918 919 920 921 922 923 924 925 926 927 928 929 930 931 932 933 934 935 936 937 938 939 940 941 942 943 944 945 946 947 948 949 950 951 952 953 954 955 956 957 958 959 960 961 962 963 964 965 966 967 968 969 970 971 972 973 974 975 976 977 978 979 980 981 982 983 984 985 986 987 988 989 990 991 992 993 994 995 996 997 998 999 1000 1001 1002 1003 1004 1005 1006 1007 1008 1009 1010 1011 1012 1013 1014 1015 1016 1017 1018 1019 1020 1021 1022 1023 1024 1025 1026 1027 1028 1029 1030 1031 1032 1033 1034 1035 1036 1037 1038 1039 1040 1041 1042 1043 1044 1045 1046 1047 1048 1049 1050 1051 1052 1053 1054 1055 1056 1057 1058 1059 1060 1061 1062 1063 1064 1065 1066 1067 1068 1069 1070 1071 1072 1073 1074 1075 1076 1077 1078 1079 1080 1081 1082 1083 1084 1085 1086 1087 1088 1089 1090 1091 1092 1093 1094 1095 1096 1097 1098 1099 1100 1101 1102 1103 1104 1105 1106 1107 1108 1109 1110 1111 1112 1113 1114 1115 1116 1117 1118 1119 1120 1121 1122 1123 1124 1125 1126 1127 1128 1129 1130 1131 1132 1133 1134 1135 1136 1137 1138 1139 1140 1141 1142 1143 1144 1145 1146 1147 1148 1149 1150 1151 1152 1153 1154 1155 1156 1157 1158 1159 1160 1161 1162 1163 1164 1165 1166 1167 1168 1169 1170 1171 1172 1173@staticmethod def add_cli_args(parser: FlexibleArgumentParser) -> FlexibleArgumentParser: """Shared CLI arguments for vLLM engine.""" # Model arguments model_kwargs = get_kwargs(ModelConfig) model_group = parser.add_argument_group( title="ModelConfig", description=ModelConfig.__doc__, ) if not ("serve" in sys.argv[1:] and "--help" in sys.argv[1:]): model_group.add_argument("--model", **model_kwargs["model"]) model_group.add_argument("--runner", **model_kwargs["runner"]) model_group.add_argument("--convert", **model_kwargs["convert"]) model_group.add_argument("--tokenizer", **model_kwargs["tokenizer"]) model_group.add_argument("--tokenizer-mode", **model_kwargs["tokenizer_mode"]) model_group.add_argument( "--trust-remote-code", **model_kwargs["trust_remote_code"] ) model_group.add_argument("--dtype", **model_kwargs["dtype"]) model_group.add_argument("--seed", **model_kwargs["seed"]) model_group.add_argument("--hf-config-path", **model_kwargs["hf_config_path"]) model_group.add_argument( "--allowed-local-media-path", **model_kwargs["allowed_local_media_path"] ) model_group.add_argument( "--allowed-media-domains", **model_kwargs["allowed_media_domains"] ) model_group.add_argument("--revision", **model_kwargs["revision"]) model_group.add_argument("--code-revision", **model_kwargs["code_revision"]) model_group.add_argument( "--tokenizer-revision", **model_kwargs["tokenizer_revision"] ) model_group.add_argument("--max-model-len", **model_kwargs["max_model_len"]) model_group.add_argument("--quantization", "-q", **model_kwargs["quantization"]) model_group.add_argument("--enforce-eager", **model_kwargs["enforce_eager"]) model_group.add_argument("--max-logprobs", **model_kwargs["max_logprobs"]) model_group.add_argument("--logprobs-mode", **model_kwargs["logprobs_mode"]) model_group.add_argument( "--disable-sliding-window", **model_kwargs["disable_sliding_window"] ) model_group.add_argument( "--disable-cascade-attn", **model_kwargs["disable_cascade_attn"] ) model_group.add_argument( "--skip-tokenizer-init", **model_kwargs["skip_tokenizer_init"] ) model_group.add_argument( "--enable-prompt-embeds", **model_kwargs["enable_prompt_embeds"] ) model_group.add_argument( "--served-model-name", **model_kwargs["served_model_name"] ) model_group.add_argument("--config-format", **model_kwargs["config_format"]) # This one is a special case because it can bool # or str. TODO: Handle this in get_kwargs model_group.add_argument( "--hf-token", type=str, nargs="?", const=True, default=model_kwargs["hf_token"]["default"], help=model_kwargs["hf_token"]["help"], ) model_group.add_argument("--hf-overrides", **model_kwargs["hf_overrides"]) model_group.add_argument("--pooler-config", **model_kwargs["pooler_config"]) model_group.add_argument( "--logits-processor-pattern", **model_kwargs["logits_processor_pattern"] ) model_group.add_argument( "--generation-config", **model_kwargs["generation_config"] ) model_group.add_argument( "--override-generation-config", **model_kwargs["override_generation_config"] ) model_group.add_argument( "--enable-sleep-mode", **model_kwargs["enable_sleep_mode"] ) model_group.add_argument("--model-impl", **model_kwargs["model_impl"]) model_group.add_argument( "--override-attention-dtype", **model_kwargs["override_attention_dtype"] ) model_group.add_argument( "--logits-processors", **model_kwargs["logits_processors"] ) model_group.add_argument( "--io-processor-plugin", **model_kwargs["io_processor_plugin"] ) # Model loading arguments load_kwargs = get_kwargs(LoadConfig) load_group = parser.add_argument_group( title="LoadConfig", description=LoadConfig.__doc__, ) load_group.add_argument("--load-format", **load_kwargs["load_format"]) load_group.add_argument("--download-dir", **load_kwargs["download_dir"]) load_group.add_argument( "--safetensors-load-strategy", **load_kwargs["safetensors_load_strategy"] ) load_group.add_argument( "--model-loader-extra-config", **load_kwargs["model_loader_extra_config"] ) load_group.add_argument("--ignore-patterns", **load_kwargs["ignore_patterns"]) load_group.add_argument("--use-tqdm-on-load", **load_kwargs["use_tqdm_on_load"]) load_group.add_argument( "--pt-load-map-location", **load_kwargs["pt_load_map_location"] ) # Attention arguments attention_kwargs = get_kwargs(AttentionConfig) attention_group = parser.add_argument_group( title="AttentionConfig", description=AttentionConfig.__doc__, ) attention_group.add_argument( "--attention-backend", **attention_kwargs["backend"] ) # Structured outputs arguments structured_outputs_kwargs = get_kwargs(StructuredOutputsConfig) structured_outputs_group = parser.add_argument_group( title="StructuredOutputsConfig", description=StructuredOutputsConfig.__doc__, ) structured_outputs_group.add_argument( "--reasoning-parser", # Choices need to be validated after parsing to include plugins **structured_outputs_kwargs["reasoning_parser"], ) structured_outputs_group.add_argument( "--reasoning-parser-plugin", **structured_outputs_kwargs["reasoning_parser_plugin"], ) # Parallel arguments parallel_kwargs = get_kwargs(ParallelConfig) parallel_group = parser.add_argument_group( title="ParallelConfig", description=ParallelConfig.__doc__, ) parallel_group.add_argument( "--distributed-executor-backend", **parallel_kwargs["distributed_executor_backend"], ) parallel_group.add_argument( "--pipeline-parallel-size", "-pp", **parallel_kwargs["pipeline_parallel_size"], ) parallel_group.add_argument("--master-addr", **parallel_kwargs["master_addr"]) parallel_group.add_argument("--master-port", **parallel_kwargs["master_port"]) parallel_group.add_argument("--nnodes", "-n", **parallel_kwargs["nnodes"]) parallel_group.add_argument("--node-rank", "-r", **parallel_kwargs["node_rank"]) parallel_group.add_argument( "--tensor-parallel-size", "-tp", **parallel_kwargs["tensor_parallel_size"] ) parallel_group.add_argument( "--decode-context-parallel-size", "-dcp", **parallel_kwargs["decode_context_parallel_size"], ) parallel_group.add_argument( "--dcp-kv-cache-interleave-size", **parallel_kwargs["dcp_kv_cache_interleave_size"], ) parallel_group.add_argument( "--cp-kv-cache-interleave-size", **parallel_kwargs["cp_kv_cache_interleave_size"], ) parallel_group.add_argument( "--prefill-context-parallel-size", "-pcp", **parallel_kwargs["prefill_context_parallel_size"], ) parallel_group.add_argument( "--data-parallel-size", "-dp", **parallel_kwargs["data_parallel_size"] ) parallel_group.add_argument( "--data-parallel-rank", "-dpn", type=int, help="Data parallel rank of this instance. " "When set, enables external load balancer mode.", ) parallel_group.add_argument( "--data-parallel-start-rank", "-dpr", type=int, help="Starting data parallel rank for secondary nodes.", ) parallel_group.add_argument( "--data-parallel-size-local", "-dpl", type=int, help="Number of data parallel replicas to run on this node.", ) parallel_group.add_argument( "--data-parallel-address", "-dpa", type=str, help="Address of data parallel cluster head-node.", ) parallel_group.add_argument( "--data-parallel-rpc-port", "-dpp", type=int, help="Port for data parallel RPC communication.", ) parallel_group.add_argument( "--data-parallel-backend", "-dpb", type=str, default="mp", help='Backend for data parallel, either "mp" or "ray".', ) parallel_group.add_argument( "--data-parallel-hybrid-lb", "-dph", **parallel_kwargs["data_parallel_hybrid_lb"], ) parallel_group.add_argument( "--data-parallel-external-lb", "-dpe", **parallel_kwargs["data_parallel_external_lb"], ) parallel_group.add_argument( "--enable-expert-parallel", **parallel_kwargs["enable_expert_parallel"] ) parallel_group.add_argument( "--all2all-backend", **parallel_kwargs["all2all_backend"] ) parallel_group.add_argument("--enable-dbo", **parallel_kwargs["enable_dbo"]) parallel_group.add_argument( "--ubatch-size", **parallel_kwargs["ubatch_size"], ) parallel_group.add_argument( "--dbo-decode-token-threshold", **parallel_kwargs["dbo_decode_token_threshold"], ) parallel_group.add_argument( "--dbo-prefill-token-threshold", **parallel_kwargs["dbo_prefill_token_threshold"], ) parallel_group.add_argument( "--disable-nccl-for-dp-synchronization", **parallel_kwargs["disable_nccl_for_dp_synchronization"], ) parallel_group.add_argument("--enable-eplb", **parallel_kwargs["enable_eplb"]) parallel_group.add_argument("--eplb-config", **parallel_kwargs["eplb_config"]) parallel_group.add_argument( "--expert-placement-strategy", **parallel_kwargs["expert_placement_strategy"], ) parallel_group.add_argument( "--max-parallel-loading-workers", **parallel_kwargs["max_parallel_loading_workers"], ) parallel_group.add_argument( "--ray-workers-use-nsight", **parallel_kwargs["ray_workers_use_nsight"] ) parallel_group.add_argument( "--disable-custom-all-reduce", **parallel_kwargs["disable_custom_all_reduce"], ) parallel_group.add_argument("--worker-cls", **parallel_kwargs["worker_cls"]) parallel_group.add_argument( "--worker-extension-cls", **parallel_kwargs["worker_extension_cls"] ) # KV cache arguments cache_kwargs = get_kwargs(CacheConfig) cache_group = parser.add_argument_group( title="CacheConfig", description=CacheConfig.__doc__, ) cache_group.add_argument("--block-size", **cache_kwargs["block_size"]) cache_group.add_argument( "--gpu-memory-utilization", **cache_kwargs["gpu_memory_utilization"] ) cache_group.add_argument( "--kv-cache-memory-bytes", **cache_kwargs["kv_cache_memory_bytes"] ) cache_group.add_argument("--swap-space", **cache_kwargs["swap_space"]) cache_group.add_argument("--kv-cache-dtype", **cache_kwargs["cache_dtype"]) cache_group.add_argument( "--num-gpu-blocks-override", **cache_kwargs["num_gpu_blocks_override"] ) cache_group.add_argument( "--enable-prefix-caching", **{ **cache_kwargs["enable_prefix_caching"], "default": None, }, ) cache_group.add_argument( "--prefix-caching-hash-algo", **cache_kwargs["prefix_caching_hash_algo"] ) cache_group.add_argument("--cpu-offload-gb", **cache_kwargs["cpu_offload_gb"]) cache_group.add_argument( "--calculate-kv-scales", **cache_kwargs["calculate_kv_scales"] ) cache_group.add_argument( "--kv-sharing-fast-prefill", **cache_kwargs["kv_sharing_fast_prefill"] ) cache_group.add_argument( "--mamba-cache-dtype", **cache_kwargs["mamba_cache_dtype"] ) cache_group.add_argument( "--mamba-ssm-cache-dtype", **cache_kwargs["mamba_ssm_cache_dtype"] ) cache_group.add_argument( "--mamba-block-size", **cache_kwargs["mamba_block_size"] ) cache_group.add_argument( "--kv-offloading-size", **cache_kwargs["kv_offloading_size"] ) cache_group.add_argument( "--kv-offloading-backend", **cache_kwargs["kv_offloading_backend"] ) # Multimodal related configs multimodal_kwargs = get_kwargs(MultiModalConfig) multimodal_group = parser.add_argument_group( title="MultiModalConfig", description=MultiModalConfig.__doc__, ) multimodal_group.add_argument( "--limit-mm-per-prompt", **multimodal_kwargs["limit_per_prompt"] ) multimodal_group.add_argument( "--enable-mm-embeds", **multimodal_kwargs["enable_mm_embeds"] ) multimodal_group.add_argument( "--media-io-kwargs", **multimodal_kwargs["media_io_kwargs"] ) multimodal_group.add_argument( "--mm-processor-kwargs", **multimodal_kwargs["mm_processor_kwargs"] ) multimodal_group.add_argument( "--mm-processor-cache-gb", **multimodal_kwargs["mm_processor_cache_gb"] ) multimodal_group.add_argument( "--mm-processor-cache-type", **multimodal_kwargs["mm_processor_cache_type"] ) multimodal_group.add_argument( "--mm-shm-cache-max-object-size-mb", **multimodal_kwargs["mm_shm_cache_max_object_size_mb"], ) multimodal_group.add_argument( "--mm-encoder-tp-mode", **multimodal_kwargs["mm_encoder_tp_mode"] ) multimodal_group.add_argument( "--mm-encoder-attn-backend", **multimodal_kwargs["mm_encoder_attn_backend"], ) multimodal_group.add_argument( "--interleave-mm-strings", **multimodal_kwargs["interleave_mm_strings"] ) multimodal_group.add_argument( "--skip-mm-profiling", **multimodal_kwargs["skip_mm_profiling"] ) multimodal_group.add_argument( "--video-pruning-rate", **multimodal_kwargs["video_pruning_rate"] ) # LoRA related configs lora_kwargs = get_kwargs(LoRAConfig) lora_group = parser.add_argument_group( title="LoRAConfig", description=LoRAConfig.__doc__, ) lora_group.add_argument( "--enable-lora", action=argparse.BooleanOptionalAction, help="If True, enable handling of LoRA adapters.", ) lora_group.add_argument("--max-loras", **lora_kwargs["max_loras"]) lora_group.add_argument("--max-lora-rank", **lora_kwargs["max_lora_rank"]) lora_group.add_argument( "--lora-dtype", **lora_kwargs["lora_dtype"], ) lora_group.add_argument("--max-cpu-loras", **lora_kwargs["max_cpu_loras"]) lora_group.add_argument( "--fully-sharded-loras", **lora_kwargs["fully_sharded_loras"] ) lora_group.add_argument("--default-mm-loras", **lora_kwargs["default_mm_loras"]) # Observability arguments observability_kwargs = get_kwargs(ObservabilityConfig) observability_group = parser.add_argument_group( title="ObservabilityConfig", description=ObservabilityConfig.__doc__, ) observability_group.add_argument( "--show-hidden-metrics-for-version", **observability_kwargs["show_hidden_metrics_for_version"], ) observability_group.add_argument( "--otlp-traces-endpoint", **observability_kwargs["otlp_traces_endpoint"] ) # TODO: generalise this special case choices = observability_kwargs["collect_detailed_traces"]["choices"] metavar = f"{{{','.join(choices)}}}" observability_kwargs["collect_detailed_traces"]["metavar"] = metavar observability_kwargs["collect_detailed_traces"]["choices"] += [ ",".join(p) for p in permutations(get_args(DetailedTraceModules), r=2) ] observability_group.add_argument( "--collect-detailed-traces", **observability_kwargs["collect_detailed_traces"], ) observability_group.add_argument( "--kv-cache-metrics", **observability_kwargs["kv_cache_metrics"] ) observability_group.add_argument( "--kv-cache-metrics-sample", **observability_kwargs["kv_cache_metrics_sample"], ) observability_group.add_argument( "--cudagraph-metrics", **observability_kwargs["cudagraph_metrics"], ) observability_group.add_argument( "--enable-layerwise-nvtx-tracing", **observability_kwargs["enable_layerwise_nvtx_tracing"], ) observability_group.add_argument( "--enable-mfu-metrics", **observability_kwargs["enable_mfu_metrics"], ) # Scheduler arguments scheduler_kwargs = get_kwargs(SchedulerConfig) scheduler_group = parser.add_argument_group( title="SchedulerConfig", description=SchedulerConfig.__doc__, ) scheduler_group.add_argument( "--max-num-batched-tokens", **{ **scheduler_kwargs["max_num_batched_tokens"], "default": None, }, ) scheduler_group.add_argument( "--max-num-seqs", **{ **scheduler_kwargs["max_num_seqs"], "default": None, }, ) scheduler_group.add_argument( "--max-num-partial-prefills", **scheduler_kwargs["max_num_partial_prefills"] ) scheduler_group.add_argument( "--max-long-partial-prefills", **scheduler_kwargs["max_long_partial_prefills"], ) scheduler_group.add_argument( "--long-prefill-token-threshold", **scheduler_kwargs["long_prefill_token_threshold"], ) # multi-step scheduling has been removed; corresponding arguments # are no longer supported. scheduler_group.add_argument( "--scheduling-policy", **scheduler_kwargs["policy"] ) scheduler_group.add_argument( "--enable-chunked-prefill", **{ **scheduler_kwargs["enable_chunked_prefill"], "default": None, }, ) scheduler_group.add_argument( "--disable-chunked-mm-input", **scheduler_kwargs["disable_chunked_mm_input"] ) scheduler_group.add_argument( "--scheduler-cls", **scheduler_kwargs["scheduler_cls"] ) scheduler_group.add_argument( "--disable-hybrid-kv-cache-manager", **scheduler_kwargs["disable_hybrid_kv_cache_manager"], ) scheduler_group.add_argument( "--async-scheduling", **scheduler_kwargs["async_scheduling"] ) scheduler_group.add_argument( "--stream-interval", **scheduler_kwargs["stream_interval"] ) # Compilation arguments compilation_kwargs = get_kwargs(CompilationConfig) compilation_group = parser.add_argument_group( title="CompilationConfig", description=CompilationConfig.__doc__, ) compilation_group.add_argument( "--cudagraph-capture-sizes", **compilation_kwargs["cudagraph_capture_sizes"] ) compilation_group.add_argument( "--max-cudagraph-capture-size", **compilation_kwargs["max_cudagraph_capture_size"], ) # vLLM arguments vllm_kwargs = get_kwargs(VllmConfig) vllm_group = parser.add_argument_group( title="VllmConfig", description=VllmConfig.__doc__, ) # We construct SpeculativeConfig using fields from other configs in # create_engine_config. So we set the type to a JSON string here to # delay the Pydantic validation that comes with SpeculativeConfig. vllm_kwargs["speculative_config"]["type"] = optional_type(json.loads) vllm_group.add_argument( "--speculative-config", **vllm_kwargs["speculative_config"] ) vllm_group.add_argument( "--kv-transfer-config", **vllm_kwargs["kv_transfer_config"] ) vllm_group.add_argument("--kv-events-config", **vllm_kwargs["kv_events_config"]) vllm_group.add_argument( "--ec-transfer-config", **vllm_kwargs["ec_transfer_config"] ) vllm_group.add_argument( "--compilation-config", "-cc", **vllm_kwargs["compilation_config"] ) vllm_group.add_argument( "--attention-config", "-ac", **vllm_kwargs["attention_config"] ) vllm_group.add_argument( "--additional-config", **vllm_kwargs["additional_config"] ) vllm_group.add_argument( "--structured-outputs-config", **vllm_kwargs["structured_outputs_config"] ) vllm_group.add_argument("--profiler-config", **vllm_kwargs["profiler_config"]) vllm_group.add_argument( "--optimization-level", **vllm_kwargs["optimization_level"] ) # Other arguments parser.add_argument( "--disable-log-stats", action="store_true", help="Disable logging statistics.", ) parser.add_argument( "--aggregate-engine-logging", action="store_true", help="Log aggregate rather than per-engine statistics " "when using data parallelism.", ) return parser create_engine_config ¶ create_engine_config( usage_context: UsageContext | None = None, headless: bool = False, ) -> VllmConfig Create the VllmConfig. NOTE: If VllmConfig is incompatible, we raise an error. Source code in vllm/engine/arg_utils.py 1312 1313 1314 1315 1316 1317 1318 1319 1320 1321 1322 1323 1324 1325 1326 1327 1328 1329 1330 1331 1332 1333 1334 1335 1336 1337 1338 1339 1340 1341 1342 1343 1344 1345 1346 1347 1348 1349 1350 1351 1352 1353 1354 1355 1356 1357 1358 1359 1360 1361 1362 1363 1364 1365 1366 1367 1368 1369 1370 1371 1372 1373 1374 1375 1376 1377 1378 1379 1380 1381 1382 1383 1384 1385 1386 1387 1388 1389 1390 1391 1392 1393 1394 1395 1396 1397 1398 1399 1400 1401 1402 1403 1404 1405 1406 1407 1408 1409 1410 1411 1412 1413 1414 1415 1416 1417 1418 1419 1420 1421 1422 1423 1424 1425 1426 1427 1428 1429 1430 1431 1432 1433 1434 1435 1436 1437 1438 1439 1440 1441 1442 1443 1444 1445 1446 1447 1448 1449 1450 1451 1452 1453 1454 1455 1456 1457 1458 1459 1460 1461 1462 1463 1464 1465 1466 1467 1468 1469 1470 1471 1472 1473 1474 1475 1476 1477 1478 1479 1480 1481 1482 1483 1484 1485 1486 1487 1488 1489 1490 1491 1492 1493 1494 1495 1496 1497 1498 1499 1500 1501 1502 1503 1504 1505 1506 1507 1508 1509 1510 1511 1512 1513 1514 1515 1516 1517 1518 1519 1520 1521 1522 1523 1524 1525 1526 1527 1528 1529 1530 1531 1532 1533 1534 1535 1536 1537 1538 1539 1540 1541 1542 1543 1544 1545 1546 1547 1548 1549 1550 1551 1552 1553 1554 1555 1556 1557 1558 1559 1560 1561 1562 1563 1564 1565 1566 1567 1568 1569 1570 1571 1572 1573 1574 1575 1576 1577 1578 1579 1580 1581 1582 1583 1584 1585 1586 1587 1588 1589 1590 1591 1592 1593 1594 1595 1596 1597 1598 1599 1600 1601 1602 1603 1604 1605 1606 1607 1608 1609 1610 1611 1612 1613 1614 1615 1616 1617 1618 1619 1620 1621 1622 1623 1624 1625 1626 1627 1628 1629 1630 1631 1632 1633 1634 1635 1636 1637 1638 1639 1640 1641 1642 1643 1644 1645 1646 1647 1648 1649 1650 1651 1652 1653 1654 1655 1656 1657 1658 1659 1660 1661 1662 1663 1664 1665 1666 1667 1668 1669 1670 1671 1672 1673 1674 1675 1676 1677 1678 1679 1680 1681 1682 1683 1684 1685 1686 1687 1688 1689 1690 1691 1692 1693 1694 1695 1696 1697 1698 1699 1700 1701 1702 1703 1704 1705 1706 1707 1708 1709 1710 1711 1712 1713 1714 1715 1716 1717 1718 1719 1720 1721 1722 1723 1724 1725 1726 1727 1728 1729 1730 1731 1732 1733 1734 1735 1736 1737def create_engine_config( self, usage_context: UsageContext | None = None, headless: bool = False, ) -> VllmConfig: """ Create the VllmConfig. NOTE: If VllmConfig is incompatible, we raise an error. """ current_platform.pre_register_and_update() device_config = DeviceConfig(device=cast(Device, current_platform.device_type)) # Check if the model is a speculator and override model/tokenizer/config # BEFORE creating ModelConfig, so the config is created with the target model # Skip speculator detection for cloud storage models (eg: S3, GCS) since # HuggingFace cannot load configs directly from S3 URLs. S3 models can still # use speculators with explicit --speculative-config. if not is_cloud_storage(self.model): (self.model, self.tokenizer, self.speculative_config) = ( maybe_override_with_speculators( model=self.model, tokenizer=self.tokenizer, revision=self.revision, trust_remote_code=self.trust_remote_code, vllm_speculative_config=self.speculative_config, ) ) model_config = self.create_model_config() self.model = model_config.model self.tokenizer = model_config.tokenizer self._check_feature_supported(model_config) self._set_default_chunked_prefill_and_prefix_caching_args(model_config) self._set_default_max_num_seqs_and_batched_tokens_args( usage_context, model_config ) sliding_window: int | None = None if not is_interleaved(model_config.hf_text_config): # Only set CacheConfig.sliding_window if the model is all sliding # window. Otherwise CacheConfig.sliding_window will override the # global layers in interleaved sliding window models. sliding_window = model_config.get_sliding_window() # Note(hc): In the current implementation of decode context # parallel(DCP), tp_size needs to be divisible by dcp_size, # because the world size does not change by dcp, it simply # reuses the GPUs of TP group, and split one TP group into # tp_size//dcp_size DCP groups. assert self.tensor_parallel_size % self.decode_context_parallel_size == 0, ( f"tp_size={self.tensor_parallel_size} must be divisible by" f"dcp_size={self.decode_context_parallel_size}." ) # Resolve "auto" kv_cache_dtype to actual value from model config resolved_cache_dtype = resolve_kv_cache_dtype_string( self.kv_cache_dtype, model_config ) cache_config = CacheConfig( block_size=self.block_size, gpu_memory_utilization=self.gpu_memory_utilization, kv_cache_memory_bytes=self.kv_cache_memory_bytes, swap_space=self.swap_space, cache_dtype=resolved_cache_dtype, is_attention_free=model_config.is_attention_free, num_gpu_blocks_override=self.num_gpu_blocks_override, sliding_window=sliding_window, enable_prefix_caching=self.enable_prefix_caching, prefix_caching_hash_algo=self.prefix_caching_hash_algo, cpu_offload_gb=self.cpu_offload_gb, calculate_kv_scales=self.calculate_kv_scales, kv_sharing_fast_prefill=self.kv_sharing_fast_prefill, mamba_cache_dtype=self.mamba_cache_dtype, mamba_ssm_cache_dtype=self.mamba_ssm_cache_dtype, mamba_block_size=self.mamba_block_size, kv_offloading_size=self.kv_offloading_size, kv_offloading_backend=self.kv_offloading_backend, ) ray_runtime_env = None if is_ray_initialized(): # Ray Serve LLM calls `create_engine_config` in the context # of a Ray task, therefore we check is_ray_initialized() # as opposed to is_in_ray_actor(). import ray ray_runtime_env = ray.get_runtime_context().runtime_env # Avoid logging sensitive environment variables sanitized_env = ray_runtime_env.to_dict() if ray_runtime_env else {} if "env_vars" in sanitized_env: sanitized_env["env_vars"] = { k: "***" for k in sanitized_env["env_vars"] } logger.info("Using ray runtime env (env vars redacted): %s", sanitized_env) # Get the current placement group if Ray is initialized and # we are in a Ray actor. If so, then the placement group will be # passed to spawned processes. placement_group = None if is_in_ray_actor(): import ray # This call initializes Ray automatically if it is not initialized, # but we should not do this here. placement_group = ray.util.get_current_placement_group() assert not headless or not self.data_parallel_hybrid_lb, ( "data_parallel_hybrid_lb is not applicable in headless mode" ) assert not (self.data_parallel_hybrid_lb and self.data_parallel_external_lb), ( "data_parallel_hybrid_lb and data_parallel_external_lb cannot both be True." ) assert self.data_parallel_backend == "mp" or self.nnodes == 1, ( "nnodes > 1 is only supported with data_parallel_backend=mp" ) inferred_data_parallel_rank = 0 if self.nnodes > 1: world_size = ( self.data_parallel_size * self.pipeline_parallel_size * self.tensor_parallel_size ) world_size_within_dp = ( self.pipeline_parallel_size * self.tensor_parallel_size ) local_world_size = world_size // self.nnodes assert world_size % self.nnodes == 0, ( f"world_size={world_size} must be divisible by nnodes={self.nnodes}." ) assert self.node_rank < self.nnodes, ( f"node_rank={self.node_rank} must be less than nnodes={self.nnodes}." ) inferred_data_parallel_rank = ( self.node_rank * local_world_size ) // world_size_within_dp if self.data_parallel_size > 1 and self.data_parallel_external_lb: self.data_parallel_rank = inferred_data_parallel_rank logger.info( "Inferred data_parallel_rank %d from node_rank %d for external lb", self.data_parallel_rank, self.node_rank, ) elif self.data_parallel_size_local is None: # Infer data parallel size local for internal dplb: self.data_parallel_size_local = max( local_world_size // world_size_within_dp, 1 ) data_parallel_external_lb = ( self.data_parallel_external_lb or self.data_parallel_rank is not None ) # Local DP rank = 1, use pure-external LB. if data_parallel_external_lb: assert self.data_parallel_rank is not None, ( "data_parallel_rank or node_rank must be specified if " "data_parallel_external_lb is enable." ) assert self.data_parallel_size_local in (1, None), ( "data_parallel_size_local must be 1 or None when data_parallel_rank " "is set" ) data_parallel_size_local = 1 # Use full external lb if we have local_size of 1. self.data_parallel_hybrid_lb = False elif self.data_parallel_size_local is not None: data_parallel_size_local = self.data_parallel_size_local if self.data_parallel_start_rank and not headless: # Infer hybrid LB mode. self.data_parallel_hybrid_lb = True if self.data_parallel_hybrid_lb and data_parallel_size_local == 1: # Use full external lb if we have local_size of 1. logger.warning( "data_parallel_hybrid_lb is not eligible when " "data_parallel_size_local = 1, autoswitch to " "data_parallel_external_lb." ) data_parallel_external_lb = True self.data_parallel_hybrid_lb = False if data_parallel_size_local == self.data_parallel_size: # Disable hybrid LB mode if set for a single node self.data_parallel_hybrid_lb = False self.data_parallel_rank = ( self.data_parallel_start_rank or inferred_data_parallel_rank ) if self.nnodes > 1: logger.info( "Inferred data_parallel_rank %d from node_rank %d", self.data_parallel_rank, self.node_rank, ) else: assert not self.data_parallel_hybrid_lb, ( "data_parallel_size_local must be set to use data_parallel_hybrid_lb." ) if self.data_parallel_backend == "ray" and ( envs.VLLM_RAY_DP_PACK_STRATEGY == "span" ): # Data parallel size defaults to 1 if DP ranks are spanning # multiple nodes data_parallel_size_local = 1 else: # Otherwise local DP size defaults to global DP size if not set data_parallel_size_local = self.data_parallel_size # DP address, used in multi-node case for torch distributed group # and ZMQ sockets. if self.data_parallel_address is None: if self.data_parallel_backend == "ray": host_ip = get_ip() logger.info( "Using host IP %s as ray-based data parallel address", host_ip ) data_parallel_address = host_ip else: assert self.data_parallel_backend == "mp", ( "data_parallel_backend can only be ray or mp, got %s", self.data_parallel_backend, ) data_parallel_address = ( self.master_addr or ParallelConfig.data_parallel_master_ip ) else: data_parallel_address = self.data_parallel_address # This port is only used when there are remote data parallel engines, # otherwise the local IPC transport is used. data_parallel_rpc_port = ( self.data_parallel_rpc_port if (self.data_parallel_rpc_port is not None) else ParallelConfig.data_parallel_rpc_port ) if self.tokens_only and not model_config.skip_tokenizer_init: model_config.skip_tokenizer_init = True logger.info("Skipping tokenizer initialization for tokens-only mode.") parallel_config = ParallelConfig( pipeline_parallel_size=self.pipeline_parallel_size, tensor_parallel_size=self.tensor_parallel_size, prefill_context_parallel_size=self.prefill_context_parallel_size, data_parallel_size=self.data_parallel_size, data_parallel_rank=self.data_parallel_rank or 0, data_parallel_external_lb=data_parallel_external_lb, data_parallel_size_local=data_parallel_size_local, master_addr=self.master_addr, master_port=self.master_port, nnodes=self.nnodes, node_rank=self.node_rank, data_parallel_master_ip=data_parallel_address, data_parallel_rpc_port=data_parallel_rpc_port, data_parallel_backend=self.data_parallel_backend, data_parallel_hybrid_lb=self.data_parallel_hybrid_lb, enable_expert_parallel=self.enable_expert_parallel, all2all_backend=self.all2all_backend, enable_dbo=self.enable_dbo, ubatch_size=self.ubatch_size, dbo_decode_token_threshold=self.dbo_decode_token_threshold, dbo_prefill_token_threshold=self.dbo_prefill_token_threshold, disable_nccl_for_dp_synchronization=self.disable_nccl_for_dp_synchronization, enable_eplb=self.enable_eplb, eplb_config=self.eplb_config, expert_placement_strategy=self.expert_placement_strategy, max_parallel_loading_workers=self.max_parallel_loading_workers, disable_custom_all_reduce=self.disable_custom_all_reduce, ray_workers_use_nsight=self.ray_workers_use_nsight, ray_runtime_env=ray_runtime_env, placement_group=placement_group, distributed_executor_backend=self.distributed_executor_backend, worker_cls=self.worker_cls, worker_extension_cls=self.worker_extension_cls, decode_context_parallel_size=self.decode_context_parallel_size, dcp_kv_cache_interleave_size=self.dcp_kv_cache_interleave_size, cp_kv_cache_interleave_size=self.cp_kv_cache_interleave_size, _api_process_count=self._api_process_count, _api_process_rank=self._api_process_rank, ) speculative_config = self.create_speculative_config( target_model_config=model_config, target_parallel_config=parallel_config, ) scheduler_config = SchedulerConfig( runner_type=model_config.runner_type, max_num_batched_tokens=self.max_num_batched_tokens, max_num_seqs=self.max_num_seqs, max_model_len=model_config.max_model_len, enable_chunked_prefill=self.enable_chunked_prefill, disable_chunked_mm_input=self.disable_chunked_mm_input, is_multimodal_model=model_config.is_multimodal_model, is_encoder_decoder=model_config.is_encoder_decoder, policy=self.scheduling_policy, scheduler_cls=self.scheduler_cls, max_num_partial_prefills=self.max_num_partial_prefills, max_long_partial_prefills=self.max_long_partial_prefills, long_prefill_token_threshold=self.long_prefill_token_threshold, disable_hybrid_kv_cache_manager=self.disable_hybrid_kv_cache_manager, async_scheduling=self.async_scheduling, stream_interval=self.stream_interval, ) if not model_config.is_multimodal_model and self.default_mm_loras: raise ValueError( "Default modality-specific LoRA(s) were provided for a " "non multimodal model" ) lora_config = ( LoRAConfig( max_lora_rank=self.max_lora_rank, max_loras=self.max_loras, default_mm_loras=self.default_mm_loras, fully_sharded_loras=self.fully_sharded_loras, lora_dtype=self.lora_dtype, max_cpu_loras=self.max_cpu_loras if self.max_cpu_loras and self.max_cpu_loras > 0 else None, ) if self.enable_lora else None ) if ( lora_config is not None and speculative_config is not None and scheduler_config.max_num_batched_tokens < ( scheduler_config.max_num_seqs * (speculative_config.num_speculative_tokens + 1) ) ): raise ValueError( "Consider increasing max_num_batched_tokens or " "decreasing num_speculative_tokens" ) # bitsandbytes pre-quantized model need a specific model loader if model_config.quantization == "bitsandbytes": self.quantization = self.load_format = "bitsandbytes" # Attention config overrides attention_config = copy.deepcopy(self.attention_config) if self.attention_backend is not None: if attention_config.backend is not None: raise ValueError( "attention_backend and attention_config.backend " "are mutually exclusive" ) # Convert string to enum if needed (CLI parsing returns a string) if isinstance(self.attention_backend, str): attention_config.backend = AttentionBackendEnum[ self.attention_backend.upper() ] else: attention_config.backend = self.attention_backend load_config = self.create_load_config() # Pass reasoning_parser into StructuredOutputsConfig if self.reasoning_parser: self.structured_outputs_config.reasoning_parser = self.reasoning_parser if self.reasoning_parser_plugin: self.structured_outputs_config.reasoning_parser_plugin = ( self.reasoning_parser_plugin ) observability_config = ObservabilityConfig( show_hidden_metrics_for_version=self.show_hidden_metrics_for_version, otlp_traces_endpoint=self.otlp_traces_endpoint, collect_detailed_traces=self.collect_detailed_traces, kv_cache_metrics=self.kv_cache_metrics, kv_cache_metrics_sample=self.kv_cache_metrics_sample, cudagraph_metrics=self.cudagraph_metrics, enable_layerwise_nvtx_tracing=self.enable_layerwise_nvtx_tracing, enable_mfu_metrics=self.enable_mfu_metrics, ) # Compilation config overrides compilation_config = copy.deepcopy(self.compilation_config) if self.cudagraph_capture_sizes is not None: if compilation_config.cudagraph_capture_sizes is not None: raise ValueError( "cudagraph_capture_sizes and compilation_config." "cudagraph_capture_sizes are mutually exclusive" ) compilation_config.cudagraph_capture_sizes = self.cudagraph_capture_sizes if self.max_cudagraph_capture_size is not None: if compilation_config.max_cudagraph_capture_size is not None: raise ValueError( "max_cudagraph_capture_size and compilation_config." "max_cudagraph_capture_size are mutually exclusive" ) compilation_config.max_cudagraph_capture_size = ( self.max_cudagraph_capture_size ) config = VllmConfig( model_config=model_config, cache_config=cache_config, parallel_config=parallel_config, scheduler_config=scheduler_config, device_config=device_config, load_config=load_config, attention_config=attention_config, lora_config=lora_config, speculative_config=speculative_config, structured_outputs_config=self.structured_outputs_config, observability_config=observability_config, compilation_config=compilation_config, kv_transfer_config=self.kv_transfer_config, kv_events_config=self.kv_events_config, ec_transfer_config=self.ec_transfer_config, profiler_config=self.profiler_config, additional_config=self.additional_config, optimization_level=self.optimization_level, ) return config create_load_config ¶ create_load_config() -> LoadConfig Source code in vllm/engine/arg_utils.py 1259 1260 1261 1262 1263 1264 1265 1266 1267 1268 1269 1270 1271 1272 1273 1274 1275 1276 1277 1278 1279 1280 1281 1282 1283def create_load_config(self) -> LoadConfig: if self.quantization == "bitsandbytes": self.load_format = "bitsandbytes" if self.load_format == "tensorizer": if hasattr(self.model_loader_extra_config, "to_serializable"): self.model_loader_extra_config = ( self.model_loader_extra_config.to_serializable() ) self.model_loader_extra_config["tensorizer_config"] = {} self.model_loader_extra_config["tensorizer_config"]["tensorizer_dir"] = ( self.model ) self.validate_tensorizer_args() return LoadConfig( load_format=self.load_format, download_dir=self.download_dir, safetensors_load_strategy=self.safetensors_load_strategy, device="cpu" if is_online_quantization(self.quantization) else None, model_loader_extra_config=self.model_loader_extra_config, ignore_patterns=self.ignore_patterns, use_tqdm_on_load=self.use_tqdm_on_load, pt_load_map_location=self.pt_load_map_location, ) create_model_config ¶ create_model_config() -> ModelConfig Source code in vllm/engine/arg_utils.py 1185 1186 1187 1188 1189 1190 1191 1192 1193 1194 1195 1196 1197 1198 1199 1200 1201 1202 1203 1204 1205 1206 1207 1208 1209 1210 1211 1212 1213 1214 1215 1216 1217 1218 1219 1220 1221 1222 1223 1224 1225 1226 1227 1228 1229 1230 1231 1232 1233 1234 1235 1236 1237 1238 1239 1240 1241 1242 1243 1244 1245 1246 1247 1248def create_model_config(self) -> ModelConfig: # gguf file needs a specific model loader if is_gguf(self.model): self.quantization = self.load_format = "gguf" if not envs.VLLM_ENABLE_V1_MULTIPROCESSING: logger.warning( "The global random seed is set to %d. Since " "VLLM_ENABLE_V1_MULTIPROCESSING is set to False, this may " "affect the random state of the Python process that " "launched vLLM.", self.seed, ) return ModelConfig( model=self.model, hf_config_path=self.hf_config_path, runner=self.runner, convert=self.convert, tokenizer=self.tokenizer, tokenizer_mode=self.tokenizer_mode, trust_remote_code=self.trust_remote_code, allowed_local_media_path=self.allowed_local_media_path, allowed_media_domains=self.allowed_media_domains, dtype=self.dtype, seed=self.seed, revision=self.revision, code_revision=self.code_revision, hf_token=self.hf_token, hf_overrides=self.hf_overrides, tokenizer_revision=self.tokenizer_revision, max_model_len=self.max_model_len, quantization=self.quantization, enforce_eager=self.enforce_eager, max_logprobs=self.max_logprobs, logprobs_mode=self.logprobs_mode, disable_sliding_window=self.disable_sliding_window, disable_cascade_attn=self.disable_cascade_attn, skip_tokenizer_init=self.skip_tokenizer_init, enable_prompt_embeds=self.enable_prompt_embeds, served_model_name=self.served_model_name, limit_mm_per_prompt=self.limit_mm_per_prompt, enable_mm_embeds=self.enable_mm_embeds, interleave_mm_strings=self.interleave_mm_strings, media_io_kwargs=self.media_io_kwargs, skip_mm_profiling=self.skip_mm_profiling, config_format=self.config_format, mm_processor_kwargs=self.mm_processor_kwargs, mm_processor_cache_gb=self.mm_processor_cache_gb, mm_processor_cache_type=self.mm_processor_cache_type, mm_shm_cache_max_object_size_mb=self.mm_shm_cache_max_object_size_mb, mm_encoder_tp_mode=self.mm_encoder_tp_mode, mm_encoder_attn_backend=self.mm_encoder_attn_backend, pooler_config=self.pooler_config, logits_processor_pattern=self.logits_processor_pattern, generation_config=self.generation_config, override_generation_config=self.override_generation_config, enable_sleep_mode=self.enable_sleep_mode, model_impl=self.model_impl, override_attention_dtype=self.override_attention_dtype, logits_processors=self.logits_processors, video_pruning_rate=self.video_pruning_rate, io_processor_plugin=self.io_processor_plugin, ) create_speculative_config ¶ create_speculative_config( target_model_config: ModelConfig, target_parallel_config: ParallelConfig, ) -> SpeculativeConfig | None Initializes and returns a SpeculativeConfig object based on speculative_config. This function utilizes speculative_config to create a SpeculativeConfig object. The speculative_config can either be provided as a JSON string input via CLI arguments or directly as a dictionary from the engine. Source code in vllm/engine/arg_utils.py 1285 1286 1287 1288 1289 1290 1291 1292 1293 1294 1295 1296 1297 1298 1299 1300 1301 1302 1303 1304 1305 1306 1307 1308 1309 1310def create_speculative_config( self, target_model_config: ModelConfig, target_parallel_config: ParallelConfig, ) -> SpeculativeConfig | None: """Initializes and returns a SpeculativeConfig object based on `speculative_config`. This function utilizes `speculative_config` to create a SpeculativeConfig object. The `speculative_config` can either be provided as a JSON string input via CLI arguments or directly as a dictionary from the engine. """ if self.speculative_config is None: return None # Note(Shangming): These parameters are not obtained from the cli arg # '--speculative-config' and must be passed in when creating the engine # config. self.speculative_config.update( { "target_model_config": target_model_config, "target_parallel_config": target_parallel_config, } ) return SpeculativeConfig(**self.speculative_config) from_cli_args classmethod ¶ from_cli_args(args: Namespace) Source code in vllm/engine/arg_utils.py 1175 1176 1177 1178 1179 1180 1181 1182 1183@classmethod def from_cli_args(cls, args: argparse.Namespace): # Get the list of attributes of this dataclass. attrs = [attr.name for attr in dataclasses.fields(cls)] # Set the attributes from the parsed arguments. engine_args = cls( **{attr: getattr(args, attr) for attr in attrs if hasattr(args, attr)} ) return engine_args get_batch_defaults classmethod ¶ get_batch_defaults( world_size: int, ) -> tuple[ dict[UsageContext | None, int], dict[UsageContext | None, int], ] Source code in vllm/engine/arg_utils.py 1784 1785 1786 1787 1788 1789 1790 1791 1792 1793 1794 1795 1796 1797 1798 1799 1800 1801 1802 1803 1804 1805 1806 1807 1808 1809 1810 1811 1812 1813 1814 1815 1816 1817 1818 1819 1820 1821 1822 1823 1824 1825 1826 1827 1828 1829 1830 1831 1832 1833 1834 1835 1836 1837 1838 1839 1840 1841 1842 1843 1844 1845 1846 1847 1848 1849 1850 1851 1852 1853 1854 1855 1856 1857 1858 1859 1860 1861 1862 1863 1864 1865 1866@classmethod def get_batch_defaults( cls, world_size: int, ) -> tuple[dict[UsageContext | None, int], dict[UsageContext | None, int]]: from vllm.usage.usage_lib import UsageContext default_max_num_batched_tokens: dict[UsageContext | None, int] default_max_num_seqs: dict[UsageContext | None, int] # When no user override, set the default values based on the usage # context. # Use different default values for different hardware. # Try to query the device name on the current platform. If it fails, # it may be because the platform that imports vLLM is not the same # as the platform that vLLM is running on (e.g. the case of scaling # vLLM with Ray) and has no GPUs. In this case we use the default # values for non-H100/H200 GPUs. try: device_memory = current_platform.get_device_total_memory() device_name = current_platform.get_device_name().lower() except Exception: # This is only used to set default_max_num_batched_tokens device_memory = 0 device_name = "" # NOTE(Kuntai): Setting large `max_num_batched_tokens` for A100 reduces # throughput, see PR #17885 for more details. # So here we do an extra device name check to prevent such regression. if device_memory >= 70 * GiB_bytes and "a100" not in device_name: # For GPUs like H100 and MI300x, use larger default values. default_max_num_batched_tokens = { UsageContext.LLM_CLASS: 16384, UsageContext.OPENAI_API_SERVER: 8192, } default_max_num_seqs = { UsageContext.LLM_CLASS: 1024, UsageContext.OPENAI_API_SERVER: 1024, } else: # TODO(woosuk): Tune the default values for other hardware. default_max_num_batched_tokens = { UsageContext.LLM_CLASS: 8192, UsageContext.OPENAI_API_SERVER: 2048, } default_max_num_seqs = { UsageContext.LLM_CLASS: 256, UsageContext.OPENAI_API_SERVER: 256, } # tpu specific default values. if current_platform.is_tpu(): chip_name = current_platform.get_device_name() if chip_name == "V6E": default_max_num_batched_tokens = { UsageContext.LLM_CLASS: 2048, UsageContext.OPENAI_API_SERVER: 1024, } elif chip_name == "V5E": default_max_num_batched_tokens = { UsageContext.LLM_CLASS: 1024, UsageContext.OPENAI_API_SERVER: 512, } elif chip_name == "V5P": default_max_num_batched_tokens = { UsageContext.LLM_CLASS: 512, UsageContext.OPENAI_API_SERVER: 256, } # cpu specific default values. if current_platform.is_cpu(): default_max_num_batched_tokens = { UsageContext.LLM_CLASS: 4096 * world_size, UsageContext.OPENAI_API_SERVER: 2048 * world_size, } default_max_num_seqs = { UsageContext.LLM_CLASS: 256 * world_size, UsageContext.OPENAI_API_SERVER: 128 * world_size, } return default_max_num_batched_tokens, default_max_num_seqs validate_tensorizer_args ¶ validate_tensorizer_args() Source code in vllm/engine/arg_utils.py 1250 1251 1252 1253 1254 1255 1256 1257def validate_tensorizer_args(self): from vllm.model_executor.model_loader.tensorizer import TensorizerConfig for key in self.model_loader_extra_config: if key in TensorizerConfig._fields: self.model_loader_extra_config["tensorizer_config"][key] = ( self.model_loader_extra_config[key] ) _compute_kwargs cached ¶ _compute_kwargs( cls: ConfigType, ) -> dict[str, dict[str, Any]] Source code in vllm/engine/arg_utils.py 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336@functools.lru_cache(maxsize=30) def _compute_kwargs(cls: ConfigType) -> dict[str, dict[str, Any]]: # Save time only getting attr docs if we're generating help text cls_docs = get_attr_docs(cls) if NEEDS_HELP else {} kwargs = {} for field in fields(cls): # Get the set of possible types for the field type_hints: set[TypeHint] = get_type_hints(field.type) # If the field is a dataclass, we can use the model_validate_json generator = (th for th in type_hints if is_dataclass(th)) dataclass_cls = next(generator, None) # Get the default value of the field if field.default is not MISSING: default = field.default # Handle pydantic.Field defaults if isinstance(default, FieldInfo): if default.default_factory is None: default = default.default else: # VllmConfig's Fields have default_factory set to config classes. # These could emit logs on init, which would be confusing. with suppress_logging(): default = default.default_factory() elif field.default_factory is not MISSING: default = field.default_factory() # Get the help text for the field name = field.name help = cls_docs.get(name, "").strip() # Escape % for argparse help = help.replace("%", "%%") # Initialise the kwargs dictionary for the field kwargs[name] = {"default": default, "help": help} # Set other kwargs based on the type hints json_tip = ( "Should either be a valid JSON string or JSON keys passed individually." ) if dataclass_cls is not None: def parse_dataclass(val: str, cls=dataclass_cls) -> Any: try: return TypeAdapter(cls).validate_json(val) except ValidationError as e: raise argparse.ArgumentTypeError(repr(e)) from e kwargs[name]["type"] = parse_dataclass kwargs[name]["help"] += f"\n\n{json_tip}" elif contains_type(type_hints, bool): # Creates --no- and -- flags kwargs[name]["action"] = argparse.BooleanOptionalAction elif contains_type(type_hints, Literal): kwargs[name].update(literal_to_kwargs(type_hints)) elif contains_type(type_hints, tuple): kwargs[name].update(collection_to_kwargs(type_hints, tuple)) elif contains_type(type_hints, list): kwargs[name].update(collection_to_kwargs(type_hints, list)) elif contains_type(type_hints, set): kwargs[name].update(collection_to_kwargs(type_hints, set)) elif contains_type(type_hints, int): if name == "max_model_len": kwargs[name]["type"] = human_readable_int_or_auto kwargs[name]["help"] += f"\n\n{human_readable_int_or_auto.__doc__}" elif name in ("max_num_batched_tokens", "kv_cache_memory_bytes"): kwargs[name]["type"] = human_readable_int kwargs[name]["help"] += f"\n\n{human_readable_int.__doc__}" else: kwargs[name]["type"] = int elif contains_type(type_hints, float): kwargs[name]["type"] = float elif contains_type(type_hints, dict) and ( contains_type(type_hints, str) or any(is_not_builtin(th) for th in type_hints) ): kwargs[name]["type"] = union_dict_and_str elif contains_type(type_hints, dict): kwargs[name]["type"] = parse_type(json.loads) kwargs[name]["help"] += f"\n\n{json_tip}" elif contains_type(type_hints, str) or any( is_not_builtin(th) for th in type_hints ): kwargs[name]["type"] = str else: raise ValueError(f"Unsupported type {type_hints} for argument {name}.") # If the type hint was a sequence of literals, use the helper function # to update the type and choices if get_origin(kwargs[name].get("type")) is Literal: kwargs[name].update(literal_to_kwargs({kwargs[name]["type"]})) # If None is in type_hints, make the argument optional. # But not if it's a bool, argparse will handle this better. if type(None) in type_hints and not contains_type(type_hints, bool): kwargs[name]["type"] = optional_type(kwargs[name]["type"]) if kwargs[name].get("choices"): kwargs[name]["choices"].append("None") return kwargs _raise_unsupported_error ¶ _raise_unsupported_error(feature_name: str) Source code in vllm/engine/arg_utils.py 2035 2036 2037 2038 2039 2040def _raise_unsupported_error(feature_name: str): msg = ( f"{feature_name} is not supported. We recommend to " f"remove {feature_name} from your config." ) raise NotImplementedError(msg) collection_to_kwargs ¶ collection_to_kwargs( type_hints: set[TypeHint], type: TypeHint ) -> dict[str, Any] Source code in vllm/engine/arg_utils.py 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200def collection_to_kwargs(type_hints: set[TypeHint], type: TypeHint) -> dict[str, Any]: type_hint = get_type(type_hints, type) types = get_args(type_hint) elem_type = types[0] # Handle Ellipsis assert all(t is elem_type for t in types if t is not Ellipsis), ( f"All non-Ellipsis elements must be of the same type. Got {types}." ) # Handle Union types if get_origin(elem_type) in {Union, UnionType}: # Union for Union[X, Y] and UnionType for X | Y assert str in get_args(elem_type), ( "If element can have multiple types, one must be 'str' " f"(i.e. 'list[int | str]'). Got {elem_type}." ) elem_type = str return { "type": elem_type, "nargs": "+" if type is not tuple or Ellipsis in types else len(types), } contains_type ¶ contains_type( type_hints: set[TypeHint], type: TypeHintT ) -> bool Check if the type hints contain a specific type. Source code in vllm/engine/arg_utils.py 151 152 153def contains_type(type_hints: set[TypeHint], type: TypeHintT) -> bool: """Check if the type hints contain a specific type.""" return any(is_type(type_hint, type) for type_hint in type_hints) get_kwargs ¶ get_kwargs(cls: ConfigType) -> dict[str, dict[str, Any]] Return argparse kwargs for the given Config dataclass. If --help or mkdocs are not present in the command line command, the attribute documentation will not be included in the help output. The heavy computation is cached via functools.lru_cache, and a deep copy is returned so callers can mutate the dictionary without affecting the cached version. Source code in vllm/engine/arg_utils.py 339 340 341 342 343 344 345 346 347 348 349def get_kwargs(cls: ConfigType) -> dict[str, dict[str, Any]]: """Return argparse kwargs for the given Config dataclass. If `--help` or `mkdocs` are not present in the command line command, the attribute documentation will not be included in the help output. The heavy computation is cached via functools.lru_cache, and a deep copy is returned so callers can mutate the dictionary without affecting the cached version. """ return copy.deepcopy(_compute_kwargs(cls)) get_type ¶ get_type( type_hints: set[TypeHint], type: TypeHintT ) -> TypeHintT Get the specific type from the type hints. Source code in vllm/engine/arg_utils.py 156 157 158def get_type(type_hints: set[TypeHint], type: TypeHintT) -> TypeHintT: """Get the specific type from the type hints.""" return next((th for th in type_hints if is_type(th, type)), None) get_type_hints ¶ get_type_hints(type_hint: TypeHint) -> set[TypeHint] Extract type hints from Annotated or Union type hints. Source code in vllm/engine/arg_utils.py 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223def get_type_hints(type_hint: TypeHint) -> set[TypeHint]: """Extract type hints from Annotated or Union type hints.""" type_hints: set[TypeHint] = set() origin = get_origin(type_hint) args = get_args(type_hint) if origin is Annotated: type_hints.update(get_type_hints(args[0])) elif origin in {Union, UnionType}: # Union for Union[X, Y] and UnionType for X | Y for arg in args: type_hints.update(get_type_hints(arg)) else: type_hints.add(type_hint) return type_hints human_readable_int ¶ human_readable_int(value: str) -> int Parse human-readable integers like '1k', '2M', etc. Including decimal values with decimal multipliers. Examples: - '1k' -> 1,000 - '1K' -> 1,024 - '25.6k' -> 25,600 Source code in vllm/engine/arg_utils.py 2043 2044 2045 2046 2047 2048 2049 2050 2051 2052 2053 2054 2055 2056 2057 2058 2059 2060 2061 2062 2063 2064 2065 2066 2067 2068 2069 2070 2071 2072 2073 2074 2075 2076 2077 2078 2079 2080 2081 2082 2083 2084 2085 2086def human_readable_int(value: str) -> int: """Parse human-readable integers like '1k', '2M', etc. Including decimal values with decimal multipliers. Examples: - '1k' -> 1,000 - '1K' -> 1,024 - '25.6k' -> 25,600 """ value = value.strip() match = re.fullmatch(r"(\d+(?:\.\d+)?)([kKmMgGtT])", value) if match: decimal_multiplier = { "k": 10**3, "m": 10**6, "g": 10**9, "t": 10**12, } binary_multiplier = { "K": 2**10, "M": 2**20, "G": 2**30, "T": 2**40, } number, suffix = match.groups() if suffix in decimal_multiplier: mult = decimal_multiplier[suffix] return int(float(number) * mult) elif suffix in binary_multiplier: mult = binary_multiplier[suffix] # Do not allow decimals with binary multipliers try: return int(number) * mult except ValueError as e: raise argparse.ArgumentTypeError( "Decimals are not allowed " f"with binary suffixes like {suffix}. Did you mean to use " f"{number}{suffix.lower()} instead?" ) from e # Regular plain number. return int(value) human_readable_int_or_auto ¶ human_readable_int_or_auto(value: str) -> int Parse human-readable integers like '1k', '2M', etc. Including decimal values with decimal multipliers. Also accepts -1 or 'auto' as a special value for auto-detection. Examples: - '1k' -> 1,000 - '1K' -> 1,024 - '25.6k' -> 25,600 - '-1' or 'auto' -> -1 (special value for auto-detection) Source code in vllm/engine/arg_utils.py 2089 2090 2091 2092 2093 2094 2095 2096 2097 2098 2099 2100 2101 2102 2103 2104 2105def human_readable_int_or_auto(value: str) -> int: """Parse human-readable integers like '1k', '2M', etc. Including decimal values with decimal multipliers. Also accepts -1 or 'auto' as a special value for auto-detection. Examples: - '1k' -> 1,000 - '1K' -> 1,024 - '25.6k' -> 25,600 - '-1' or 'auto' -> -1 (special value for auto-detection) """ value = value.strip() if value == "-1" or value.lower() == "auto": return -1 return human_readable_int(value) is_not_builtin ¶ is_not_builtin(type_hint: TypeHint) -> bool Check if the class is not a built-in type. Source code in vllm/engine/arg_utils.py 203 204 205def is_not_builtin(type_hint: TypeHint) -> bool: """Check if the class is not a built-in type.""" return type_hint.__module__ != "builtins" is_online_quantization ¶ is_online_quantization(quantization: Any) -> bool Source code in vllm/engine/arg_utils.py 226 227def is_online_quantization(quantization: Any) -> bool: return quantization in ["inc"] is_type ¶ is_type( type_hint: TypeHint, type: TypeHintT ) -> TypeIs[TypeHintT] Check if the type hint is a specific type. Source code in vllm/engine/arg_utils.py 146 147 148def is_type(type_hint: TypeHint, type: TypeHintT) -> TypeIs[TypeHintT]: """Check if the type hint is a specific type.""" return type_hint is type or get_origin(type_hint) is type literal_to_kwargs ¶ literal_to_kwargs( type_hints: set[TypeHint], ) -> dict[str, Any] Get the type and choices from a Literal type hint in type_hints. If type_hints also contains str, we use metavar instead of choices. Source code in vllm/engine/arg_utils.py 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175def literal_to_kwargs(type_hints: set[TypeHint]) -> dict[str, Any]: """Get the `type` and `choices` from a `Literal` type hint in `type_hints`. If `type_hints` also contains `str`, we use `metavar` instead of `choices`. """ type_hint = get_type(type_hints, Literal) options = get_args(type_hint) option_type = type(options[0]) if not all(isinstance(option, option_type) for option in options): raise ValueError( "All options must be of the same type. " f"Got {options} with types {[type(c) for c in options]}" ) kwarg = "metavar" if contains_type(type_hints, str) else "choices" return {"type": option_type, kwarg: sorted(options)} optional_type ¶ optional_type( return_type: Callable[[str], T], ) -> Callable[[str], T | None] Source code in vllm/engine/arg_utils.py 131 132 133 134 135 136 137def optional_type(return_type: Callable[[str], T]) -> Callable[[str], T | None]: def _optional_type(val: str) -> T | None: if val == "" or val == "None": return None return parse_type(return_type)(val) return _optional_type parse_type ¶ parse_type( return_type: Callable[[str], T], ) -> Callable[[str], T] Source code in vllm/engine/arg_utils.py 119 120 121 122 123 124 125 126 127 128def parse_type(return_type: Callable[[str], T]) -> Callable[[str], T]: def _parse_type(val: str) -> T: try: return return_type(val) except ValueError as e: raise argparse.ArgumentTypeError( f"Value {val} cannot be converted to {return_type}." ) from e return _parse_type union_dict_and_str ¶ union_dict_and_str(val: str) -> str | dict[str, str] | None Source code in vllm/engine/arg_utils.py 140 141 142 143def union_dict_and_str(val: str) -> str | dict[str, str] | None: if not re.match(r"(?s)^\s*{.*}\s*$", val): return str(val) return optional_type(json.loads)(val) December 25, 2025 ``` ``` **Pattern 8:** vLLM GitHub Home User Guide User Guide Getting Started Getting Started Quickstart Installation Installation GPU CPU TPU Examples Examples Offline inference Offline inference Async LLM Streaming Audio Language Automatic Prefix Caching Basic Batch LLM Inference Chat With Tools Context Extension Data Parallel Disaggregated Prefill V1 Disaggregated Prefill Encoder Decoder Multimodal KV Load Failure Recovery Test LLM Engine Example LLM Engine Reset Kv Load Sharded State Logits Processor LoRA With Quantization Inference Metrics Mistral-Small MLPSpeculator MultiLoRA Inference Offline Inference with the OpenAI Batch file format Prefix Caching Prompt Embed Inference Qwen2.5-Omni Offline Inference Examples Qwen3 Omni Qwen 1M Reproducibility RLHF RLHF Colocate RLHF Online Quant RLHF Utils Save Sharded State Simple Profiling Skip Loading Weights In Engine Init Spec Decode Structured Outputs Torchrun Dp Example Torchrun Example Vision Language Vision Language Multi Image Online serving Online serving API Client Helm Charts Monitoring Dashboards Disaggregated Encoder Disaggregated Prefill Disaggregated Serving Disaggregated Serving P2P Nccl Xpyd Elastic Ep Gradio OpenAI Chatbot Webserver Gradio Webserver Kv Events Subscriber Multi-Node-Serving Multi Instance Data Parallel OpenAI Chat Completion Client OpenAI Chat Completion Client For Multimodal OpenAI Chat Completion Client With Tools OpenAI Chat Completion Client With Tools Required OpenAI Chat Completion Client With Tools Xlam OpenAI Chat Completion Client With Tools Xlam Streaming OpenAI Chat Completion Tool Calls With Reasoning OpenAI Chat Completion With Reasoning OpenAI Chat Completion With Reasoning Streaming OpenAI Completion Client OpenAI Responses Client OpenAI Responses Client With Mcp Tools OpenAI Responses Client With Tools OpenAI Transcription Client OpenAI Translation Client Setup OpenTelemetry POC Prometheus and Grafana Prompt Embed Inference With OpenAI Client Ray Serve Deepseek Retrieval Augmented Generation With Langchain Retrieval Augmented Generation With Llamaindex Run Cluster Sagemaker-Entrypoint Streamlit OpenAI Chatbot Webserver Structured Outputs Token Generation Client Utils Others Others LMCache Examples Logging Configuration Tensorize vLLM Model Pooling Pooling Classify Embed Plugin Pooling Score Token Classify Token Embed General General vLLM V1 Frequently Asked Questions Production Metrics Reproducibility Security Troubleshooting Usage Stats Collection Inference and Serving Inference and Serving Offline Inference OpenAI-Compatible Server Context Parallel Deployment Data Parallel Deployment Troubleshooting distributed deployments Expert Parallel Deployment Parallelism and Scaling Integrations Integrations LangChain LlamaIndex Deployment Deployment Using Docker Using Kubernetes Using Nginx Frameworks Frameworks Anyscale AnythingLLM AutoGen BentoML Cerebrium Chatbox Dify dstack Haystack Helm Hugging Face Inference Endpoints LiteLLM Lobe Chat LWS Modal Open WebUI Retrieval-Augmented Generation SkyPilot Streamlit NVIDIA Triton Integrations Integrations KAITO KServe Kthena KubeAI KubeRay Llama Stack llm-d llmaz Production stack Training Training Reinforcement Learning from Human Feedback Transformers Reinforcement Learning Configuration Configuration Conserving Memory Engine Arguments Environment Variables Model Resolution Optimization and Tuning Server Arguments TPU Models Models Supported Models Generative Models Pooling Models Extensions Extensions Loading Model weights with fastsafetensors Loading models with Run:ai Model Streamer Loading models with CoreWeave's Tensorizer Hardware Supported Models Hardware Supported Models CPU - Intel® Xeon® XPU - Intel® GPUs TPU Features Features Automatic Prefix Caching Automatic Prefix Caching Table of contents Introduction Enabling APC in vLLM Example workloads Limits Batch Invariance Custom Arguments Custom Logits Processors Disaggregated Encoder Disaggregated Prefilling (experimental) Interleaved Thinking LoRA Adapters MooncakeConnector Usage Guide Multimodal Inputs NixlConnector Usage Guide Prompt Embedding Inputs Reasoning Outputs Sleep Mode Speculative Decoding Structured Outputs Tool Calling Quantization Quantization AutoAWQ AutoRound BitBLAS BitsAndBytes FP8 W8A8 GGUF GPTQModel FP8 INC INT4 W4A16 INT8 W8A8 NVIDIA Model Optimizer Quantized KV Cache AMD Quark TorchAO Developer Guide Developer Guide General General Deprecation Policy Dockerfile Incremental Compilation Workflow Profiling vLLM Vulnerability Management Model Implementation Model Implementation Basic Model Registering a Model Unit Testing Multi-Modal Support Speech-to-Text (Transcription/Translation) Support CI CI CI Failures Nightly Builds of vLLM Wheels Update PyTorch version on vLLM OSS CI/CD Design Documents Design Documents Plugins Plugins IO Processor Plugins LoRA Resolver Plugins Plugin System Architecture Overview CUDA Graphs Dual Batch Overlap How to debug the vLLM-torch.compile integration Fused MoE Modular Kernel Integration with Hugging Face Hybrid KV Cache Manager Logits Processors Metrics Multi-Modal Data Processing Fused MoE Kernel Features Python Multiprocessing Optimization levels P2P NCCL Connector Paged Attention Automatic Prefix Caching torch.compile integration Benchmarking Benchmarking Benchmark CLI Parameter Sweeps Performance Dashboard API Reference API Reference vllm vllm beam_search collect_env connections env_override envs forward_context logger logits_process logprobs outputs pooling_params sampling_params scalar_type scripts sequence tasks tracing version assets assets audio base image video attention attention layer selector backends backends abstract registry utils layers layers chunked_local_attention cross_attention encoder_only_attention mm_encoder_attention ops ops chunked_prefill_paged_decode common flashmla merge_attn_states paged_attn pallas_kv_cache_update prefix_prefill rocm_aiter_mla_sparse triton_decode_attention triton_merge_attn_states triton_reshape_and_cache_flash triton_unified_attention vit_attn_wrappers utils utils fa_utils kv_sharing_utils kv_transfer_utils benchmarks benchmarks datasets latency serve startup throughput lib lib endpoint_request_func ready_checker utils sweep sweep cli param_sweep plot plot_pareto serve serve_sla server sla_sweep utils compilation compilation activation_quant_fusion backends base_static_graph caching collective_fusion compiler_interface counter cuda_graph decorators fix_functionalization fusion fusion_attn fx_utils inductor_pass matcher_utils monitor noop_elimination partition_rules pass_manager piecewise_backend post_cleanup qk_norm_rope_fusion rocm_aiter_fusion sequence_parallelism torch25_custom_graph_pass vllm_inductor_pass wrapper config config attention cache compilation device ec_transfer kv_events kv_transfer load lora model multimodal observability parallel pooler profiler scheduler speculative speech_to_text structured_outputs utils vllm device_allocator device_allocator cumem distributed distributed communication_op kv_events parallel_state tpu_distributed_utils utils device_communicators device_communicators all2all all_reduce_utils base_device_communicator cpu_communicator cuda_communicator cuda_wrapper custom_all_reduce mnnvl_compat pynccl pynccl_allocator pynccl_wrapper quick_all_reduce ray_communicator shm_broadcast shm_object_storage symm_mem tpu_communicator xpu_communicator ec_transfer ec_transfer ec_transfer_state ec_connector ec_connector base example_connector factory eplb eplb async_worker eplb_state rebalance_execute policy policy abstract default kv_transfer kv_transfer kv_transfer_state kv_connector kv_connector base factory utils v1 v1 base decode_bench_connector example_connector lmcache_connector lmcache_mp_connector metrics mooncake_connector multi_connector nixl_connector offloading_connector lmcache_integration lmcache_integration multi_process_adapter utils vllm_v1_adapter p2p p2p p2p_nccl_connector p2p_nccl_engine tensor_memory_pool engine engine arg_utils async_llm_engine llm_engine protocol entrypoints entrypoints api_server chat_utils constants context launcher llm logger renderer responses_utils score_utils ssl tool tool_server utils anthropic anthropic protocol serving_messages cli cli collect_env main openai run_batch serve types benchmark benchmark base latency main serve startup sweep throughput openai openai api_server cli_args orca_metrics protocol run_batch serving_chat serving_chat_stream_harmony serving_completion serving_engine serving_models serving_responses serving_transcription speech_to_text utils parser parser harmony_utils responses_parser pooling pooling classify classify api_router protocol serving embed embed api_router conftest protocol serving pooling pooling api_router protocol serving score score api_router protocol serving sagemaker sagemaker routes serve serve cache cache api_router disagg disagg api_router protocol serving elastic_ep elastic_ep api_router middleware instrumentator instrumentator health metrics server_info lora lora api_router profile profile api_router rlhf rlhf api_router rpc rpc api_router sleep sleep api_router tokenize tokenize api_router serving inputs inputs data parse preprocess logging_utils logging_utils dump_input formatter lazy log_time lora lora lora_model lora_weights model_manager peft_helper request resolver utils worker_manager layers layers base base_linear column_parallel_linear fused_moe logits_processor replicated_linear row_parallel_linear utils vocal_parallel_embedding ops ops ipex_ops ipex_ops lora_ops torch_ops torch_ops lora_ops triton_ops triton_ops fused_moe_lora_op kernel_utils lora_expand_op lora_kernel_metadata lora_shrink_op utils xla_ops xla_ops lora_ops punica_wrapper punica_wrapper punica_base punica_cpu punica_gpu punica_selector punica_tpu punica_xpu utils model_executor model_executor custom_op parameter utils layers layers activation attention_layer_base batch_invariant conv kda layernorm lightning_attn linear logits_processor mla pooler resampler utils vocab_parallel_embedding fla fla ops ops chunk chunk_delta_h chunk_o chunk_scaled_dot_kkt cumsum fused_recurrent index kda l2norm layernorm_guard op solve_tril utils wy_fast fused_moe fused_moe all2all_utils batched_deep_gemm_moe config cpu_fused_moe cutlass_moe deep_gemm_moe deep_gemm_utils deepep_ht_prepare_finalize deepep_ll_prepare_finalize flashinfer_cutedsl_moe flashinfer_cutlass_moe flashinfer_cutlass_prepare_finalize flashinfer_trtllm_moe fused_batched_moe fused_marlin_moe fused_moe fused_moe_method_base fused_moe_modular_method gpt_oss_triton_kernels_moe layer modular_kernel moe_align_block_size moe_pallas moe_permute_unpermute moe_torch_iterative pplx_prepare_finalize prepare_finalize rocm_aiter_fused_moe routing_simulator shared_fused_moe topk_weight_and_reduce triton_deep_gemm_moe trtllm_moe unquantized_fused_moe_method utils zero_expert_fused_moe mamba mamba abstract linear_attn mamba_mixer mamba_mixer2 mamba_utils short_conv ops ops causal_conv1d layernorm_gated mamba_ssm ssd_bmm ssd_chunk_scan ssd_chunk_state ssd_combined ssd_state_passing quantization quantization auto_round awq awq_marlin awq_triton base_config bitblas bitsandbytes cpu_wna16 deepspeedfp experts_int8 fbgemm_fp8 fp8 fp_quant gguf gptq gptq_bitblas gptq_marlin gptq_marlin_24 hqq_marlin inc input_quant_fp8 ipex_quant kv_cache modelopt moe_wna16 mxfp4 petit ptpc_fp8 qutlass_utils rtn schema torchao tpu_int8 compressed_tensors compressed_tensors compressed_tensors compressed_tensors_moe triton_scaled_mm utils schemes schemes compressed_tensors_24 compressed_tensors_scheme compressed_tensors_w4a4_nvfp4 compressed_tensors_w4a8_fp8 compressed_tensors_w4a8_int compressed_tensors_w4a16_24 compressed_tensors_w4a16_nvfp4 compressed_tensors_w8a8_fp8 compressed_tensors_w8a8_int8 compressed_tensors_w8a16_fp8 compressed_tensors_wNa16 transform transform linear module utils schemes schemes linear_qutlass_nvfp4 kernels kernels mixed_precision mixed_precision allspark bitblas conch cutlass dynamic_4bit exllama MPLinearKernel machete marlin xpu scaled_mm scaled_mm aiter cpu cutlass ScaledMMLinearKernel triton xla quark quark quark quark_moe utils schemes schemes quark_ocp_mx quark_scheme quark_w8a8_fp8 quark_w8a8_int8 utils utils allspark_utils bitblas_utils flashinfer_fp4_moe flashinfer_utils fp8_utils gptq_utils int8_utils layer_utils machete_utils marlin_utils marlin_utils_fp4 marlin_utils_fp8 marlin_utils_test marlin_utils_test_24 mxfp4_utils mxfp6_utils mxfp8_utils nvfp4_emulation_utils nvfp4_moe_support ocp_mx_utils petit_utils quant_utils w8a8_utils rotary_embedding rotary_embedding base common deepseek_scaling_rope dual_chunk_rope dynamic_ntk_alpha_rope dynamic_ntk_scaling_rope ernie45_vl_rope linear_scaling_rope llama3_rope llama4_vision_rope mrope ntk_scaling_rope phi3_long_rope_scaled_rope xdrope yarn_scaling_rope model_loader model_loader base_loader bitsandbytes_loader default_loader dummy_loader gguf_loader online_quantization runai_streamer_loader sharded_state_loader tensorizer tensorizer_loader tpu utils weight_utils models models adapters afmoe aimv2 apertus arcee arctic aria audioflamingo3 aya_vision bagel baichuan bailing_moe bamba bee bert bert_with_rope blip blip2 bloom chameleon chatglm clip cohere2_vision commandr config dbrx deepencoder deepseek_eagle deepseek_mtp deepseek_ocr deepseek_v2 deepseek_vl2 dots1 dots_ocr ernie45 ernie45_moe ernie45_vl ernie45_vl_moe ernie_mtp exaone exaone4 fairseq2_llama falcon falcon_h1 flex_olmo fuyu gemma gemma2 gemma3 gemma3_mm gemma3n gemma3n_mm glm glm4 glm4_1v glm4_moe glm4_moe_mtp glm4v gpt2 gpt_bigcode gpt_j gpt_neox gpt_oss granite granite_speech granitemoe granitemoehybrid granitemoeshared gritlm grok1 h2ovl hunyuan_v1 hunyuan_vision hyperclovax_vision idefics2_vision_model idefics3 interfaces interfaces_base intern_vit internlm2 internlm2_ve interns1 interns1_vit internvl jais jais2 jamba jina_vl keye keye_vl1_5 kimi_linear kimi_vl lfm2 lfm2_moe lightonocr llama llama4 llama4_eagle llama_eagle llama_eagle3 llava llava_next llava_next_video llava_onevision longcat_flash longcat_flash_mtp mamba mamba2 medusa midashenglm mimo mimo_mtp mimo_v2_flash minicpm minicpm3 minicpm_eagle minicpmo minicpmv minimax_m2 minimax_text_01 minimax_vl_01 mistral3 mistral_large_3 mistral_large_3_eagle mixtral mllama4 mlp_speculator modernbert module_mapping molmo moonvit mpt nano_nemotron_vl nemotron nemotron_h nemotron_nas nemotron_vl nvlm_d olmo olmo2 olmoe opencua openpangu openpangu_mtp opt orion ouro ovis ovis2_5 paddleocr_vl paligemma persimmon phi phi3 phi3v phi4mm phi4mm_audio phi4mm_utils phimoe pixtral plamo2 plamo3 qwen qwen2 qwen2_5_omni_thinker qwen2_5_vl qwen2_audio qwen2_moe qwen2_rm qwen2_vl qwen3 qwen3_moe qwen3_next qwen3_next_mtp qwen3_omni_moe_thinker qwen3_vl qwen3_vl_moe qwen_vl radio registry roberta rvl seed_oss siglip siglip2navit skyworkr1v smolvlm solar stablelm starcoder2 step3_text step3_vl swin tarsier telechat2 teleflm terratorch ultravox utils vision voxtral voxtral_streaming whisper whisper_utils zamba2 transformers transformers base causal legacy moe multimodal pooling utils warmup warmup deep_gemm_warmup kernel_warmup multimodal multimodal audio base cache evs hasher image inputs parse processing profiling registry utils video platforms platforms cpu cuda interface rocm tpu xpu plugins plugins io_processors io_processors interface lora_resolvers lora_resolvers filesystem_resolver profiler profiler layerwise_profile utils wrapper ray ray lazy_utils ray_env reasoning reasoning abs_reasoning_parsers basic_parsers deepseek_r1_reasoning_parser deepseek_v3_reasoning_parser ernie45_reasoning_parser glm4_moe_reasoning_parser gptoss_reasoning_parser granite_reasoning_parser holo2_reasoning_parser hunyuan_a13b_reasoning_parser identity_reasoning_parser minimax_m2_reasoning_parser mistral_reasoning_parser olmo3_reasoning_parser qwen3_reasoning_parser seedoss_reasoning_parser step3_reasoning_parser tokenizers tokenizers deepseek_v32 deepseek_v32_encoding detokenizer_utils hf mistral protocol registry tool_parsers tool_parsers abstract_tool_parser deepseekv3_tool_parser deepseekv31_tool_parser deepseekv32_tool_parser ernie45_tool_parser functiongemma_tool_parser gigachat3_tool_parser glm4_moe_tool_parser glm47_moe_tool_parser granite_20b_fc_tool_parser granite_tool_parser hermes_tool_parser hunyuan_a13b_tool_parser internlm2_tool_parser jamba_tool_parser kimi_k2_tool_parser llama4_pythonic_tool_parser llama_tool_parser longcat_tool_parser minimax_m2_tool_parser minimax_tool_parser mistral_tool_parser olmo3_tool_parser openai_tool_parser phi4mini_tool_parser pythonic_tool_parser qwen3coder_tool_parser qwen3xml_tool_parser seed_oss_tool_parser step3_tool_parser utils xlam_tool_parser transformers_utils transformers_utils config config_parser_base dynamic_module gguf_utils processor repo_utils runai_utils s3_utils tokenizer utils chat_templates chat_templates registry configs configs afmoe arctic bagel chatglm deepseek_vl2 dotsocr eagle falcon flex_olmo hunyuan_vl jais kimi_linear kimi_vl lfm2_moe medusa midashenglm mistral mlp_speculator moonvit nemotron nemotron_h olmo3 ovis qwen3_next radio step3_vl tarsier2 ultravox speculators speculators algos base processors processors bagel deepseek_ocr deepseek_vl2 hunyuan_vl hunyuan_vl_image ovis ovis2_5 triton_utils triton_utils importing usage usage usage_lib utils utils argparse_utils async_utils cache collection_utils counter deep_gemm flashinfer func_utils gc_utils hashing import_utils jsontree math_utils mem_constants mem_utils nccl network_utils nvtx_pytorch_hooks platform_utils profiling registry serial_utils system_utils tensor_schema torch_utils v1 v1 cudagraph_dispatcher kv_cache_interface outputs request serial_utils utils attention attention backends backends cpu_attn flash_attn flashinfer flex_attention gdn_attn linear_attn mamba1_attn mamba2_attn mamba_attn pallas rocm_aiter_fa rocm_aiter_unified_attn rocm_attn short_conv_attn tree_attn triton_attn utils mla mla aiter_triton_mla common cutlass_mla flashattn_mla flashinfer_mla flashmla flashmla_sparse indexer rocm_aiter_mla rocm_aiter_mla_sparse triton_mla core core block_pool encoder_cache_manager kv_cache_coordinator kv_cache_manager kv_cache_metrics kv_cache_utils single_type_kv_cache_manager sched sched async_scheduler interface output request_queue scheduler utils engine engine async_llm coordinator core core_client detokenizer exceptions input_processor llm_engine logprobs output_processor parallel_sampling utils executor executor abstract multiproc_executor ray_distributed_executor ray_executor ray_utils uniproc_executor kv_offload kv_offload abstract arc_manager backend cpu factory lru_manager mediums spec backends backends cpu worker worker cpu_gpu worker metrics metrics loggers perf prometheus ray_wrappers reader stats pool pool metadata sample sample metadata rejection_sampler sampler logits_processor logits_processor builtin interface state ops ops bad_words logprobs penalties topk_topp_sampler tpu tpu metadata sampler spec_decode spec_decode eagle medusa metadata metrics ngram_proposer suffix_decoding utils structured_output structured_output backend_guidance backend_lm_format_enforcer backend_outlines backend_types backend_xgrammar request utils worker worker block_table cp_utils cpu_model_runner cpu_worker dp_utils ec_connector_model_runner_mixin gpu_input_batch gpu_model_runner gpu_ubatch_wrapper gpu_worker kv_connector_model_runner_mixin lora_model_runner_mixin tpu_input_batch tpu_model_runner tpu_worker ubatch_utils ubatching utils worker_base workspace xpu_model_runner xpu_worker gpu gpu async_utils attn_utils block_table cudagraph_utils dp_utils input_batch model_runner states structured_outputs metrics metrics logits sample sample gumbel logprob metadata min_p output penalties sampler spec_decode spec_decode eagle eagle_cudagraph rejection_sample CLI Reference CLI Reference vllm serve vllm chat vllm complete vllm run-batch vllm bench vllm bench vllm bench latency vllm bench serve vllm bench sweep plot vllm bench sweep plot_pareto vllm bench sweep serve vllm bench sweep serve_sla vllm bench throughput Community Community Contact Us Meetups Sponsors Governance Governance Collaboration Policy Committers Governance Process Blog Forum Slack Table of contents Introduction Enabling APC in vLLM Example workloads Limits Automatic Prefix Caching¶ Introduction¶ Automatic Prefix Caching (APC in short) caches the KV cache of existing queries, so that a new query can directly reuse the KV cache if it shares the same prefix with one of the existing queries, allowing the new query to skip the computation of the shared part. Note Technical details on how vLLM implements APC can be found here. Enabling APC in vLLM¶ Set enable_prefix_caching=True in vLLM engine to enable APC. Here is an example: examples/offline_inference/automatic_prefix_caching.py Example workloads¶ We describe two example workloads, where APC can provide huge performance benefit: Long document query, where the user repeatedly queries the same long document (e.g. software manual or annual report) with different queries. In this case, instead of processing the long document again and again, APC allows vLLM to process this long document only once, and all future requests can avoid recomputing this long document by reusing its KV cache. This allows vLLM to serve future requests with much higher throughput and much lower latency. Multi-round conversation, where the user may chat with the application multiple times in the same chatting session. In this case, instead of processing the whole chatting history again and again, APC allows vLLM to reuse the processing results of the chat history across all future rounds of conversation, allowing vLLM to serve future requests with much higher throughput and much lower latency. Limits¶ APC in general does not reduce the performance of vLLM. With that being said, APC only reduces the time of processing the queries (the prefilling phase) and does not reduce the time of generating new tokens (the decoding phase). So APC does not bring performance gain when vLLM spends most of the time generating answers to the queries (e.g. when the length of the answer is long), or new queries do not share the same prefix with any of existing queries (so that the computation cannot be reused). October 17, 2025 ``` enable_prefix_caching=True ``` ### Example Code Patterns **Example 1** (json): ```json [CI Failure]: failing-test-job - regex/matching/failing:test ``` **Example 2** (json): ```json [CI Failure]: failing-test-job - regex/matching/failing:test ``` **Example 3** (typescript): ```typescript class BatchDescriptor(NamedTuple): num_tokens: int num_reqs: int uniform: bool = False has_lora: bool = False ``` **Example 4** (typescript): ```typescript class BatchDescriptor(NamedTuple): num_tokens: int num_reqs: int uniform: bool = False has_lora: bool = False ``` ## Reference Files This skill includes comprehensive documentation in `references/`: - **api.md** - Api documentation - **configuration.md** - Configuration documentation - **deployment.md** - Deployment documentation - **developer.md** - Developer documentation - **examples.md** - Examples documentation - **features.md** - Features documentation - **getting_started.md** - Getting Started documentation - **models.md** - Models documentation - **other.md** - Other documentation - **performance.md** - Performance documentation Use `view` to read specific reference files when detailed information is needed. ## Working with This Skill ### For Beginners Start with the getting_started or tutorials reference files for foundational concepts. ### For Specific Features Use the appropriate category reference file (api, guides, etc.) for detailed information. ### For Code Examples The quick reference section above contains common patterns extracted from the official docs. ## Resources ### references/ Organized documentation extracted from official sources. These files contain: - Detailed explanations - Code examples with language annotations - Links to original documentation - Table of contents for quick navigation ### scripts/ Add helper scripts here for common automation tasks. ### assets/ Add templates, boilerplate, or example projects here. ## Notes - This skill was automatically generated from official documentation - Reference files preserve the structure and examples from source docs - Code examples include language detection for better syntax highlighting - Quick reference patterns are extracted from common usage examples in the docs ## Updating To refresh this skill with updated documentation: 1. Re-run the scraper with the same configuration 2. The skill will be rebuilt with the latest information