vllm_mlx.patches.glm4v_moe_mllm¶
Runtime patch for mlx-vlm's GLM-4.6V model to support BatchKVCache.
View the complete module source at #L1-L89.
API details¶
Each callable below includes its exact signature, type annotations, inputs, defaults, return contract, documented exceptions, implementation source, and parsed docstring sections when the source provides them.
vllm_mlx.patches.glm4v_moe_mllm
¶
Runtime patch for mlx-vlm's GLM-4.6V model to support BatchKVCache.
GLM-4.6V (glm4v_moe) computes position_ids from cache[0].offset once at the start of GLM4VModel.call, then derives position_embeddings used by ALL decoder layers:
position_ids = mx.arange(cache[0].offset, cache[0].offset + seq_len)
position_embeddings = self.rotary_emb(h, position_ids)
for layer in self.layers:
h = layer(h, mask, cache[i], position_embeddings)
For regular KVCache, cache.offset is a Python int, so mx.arange works fine. For BatchKVCache, cache.offset is an mx.array (per-batch-item offsets), and mx.arange does not support mx.array start/stop arguments, producing wrong position_ids that corrupt RoPE embeddings for ALL layers.
This patch replaces GLM4VModel.call with a version that converts cache[0].offset to int before computing position_ids.
vllm_mlx.patches.glm4v_moe_mllm.patch_glm4v_moe_for_batching
¶
Monkey-patch GLM4VModel.call to handle BatchKVCache offset.
Returns True if patch was applied, False if mlx-vlm is not installed or GLM-4.6V module not available.
Source code in vllm_mlx/patches/glm4v_moe_mllm.py
Complete contract reference¶
Expand any definition for its exact inputs, annotations, defaults, return contract, directly raised exceptions, source-grounded behavior, and immutable line link. This section includes private and nested definitions that ordinary API generators omit.
vllm_mlx.patches.glm4v_moe_mllm.patch_glm4v_moe_for_batching · function
Monkey-patch GLM4VModel.call to handle BatchKVCache offset.
Parameters
This callable has no explicit inputs.
Returns
- Type:
bool - Direct return expressions:
False;True
Exceptions and behavior
Function patch_glm4v_moe_for_batching calls logger.debug, getattr, logger.info; has 2 explicit return paths.
No direct raise statement appears in this definition.
vllm_mlx.patches.glm4v_moe_mllm.patch_glm4v_moe_for_batching._patched_call · nested function
vllm_mlx.patches.glm4v_moe_mllm.patch_glm4v_moe_for_batching._patched_call(inputs: mx.array, inputs_embeds: Optional[mx.array] = None, cache: Optional[Any] = None, mask: Optional[mx.array] = None, position_ids: Optional[mx.array] = None) -> mx.array
Nested Function patch_glm4v_moe_for_batching._patched_call calls self.embed_tokens, inputs_embeds.astype, isinstance, int; returns self.norm(h).
Parameters
| Name | Type | Required | Default | Description |
|---|---|---|---|---|
inputs |
mx.array |
yes |
none |
Required positional or keyword input. |
inputs_embeds |
Optional[mx.array] |
no |
None |
Optional positional or keyword input; defaults to None. |
cache |
Optional[Any] |
no |
None |
Optional positional or keyword input; defaults to None. |
mask |
Optional[mx.array] |
no |
None |
Optional positional or keyword input; defaults to None. |
position_ids |
Optional[mx.array] |
no |
None |
Optional positional or keyword input; defaults to None. |
Returns
- Type:
mx.array - Direct return expressions:
self.norm(h)
Exceptions and behavior
Nested Function patch_glm4v_moe_for_batching._patched_call calls self.embed_tokens, inputs_embeds.astype, isinstance, int; returns self.norm(h).
No direct raise statement appears in this definition.
Complete symbol map¶
This map also includes private definitions and nested helpers. The signature column exposes every explicit input even when an internal helper has no dedicated parameter prose.
| Symbol | Kind | Signature and inputs | What it does | Source |
|---|---|---|---|---|
patch_glm4v_moe_for_batching |
function | patch_glm4v_moe_for_batching() -> bool |
Monkey-patch GLM4VModel.call to handle BatchKVCache offset. | #L31-L89 |
patch_glm4v_moe_for_batching._patched_call |
nested function | patch_glm4v_moe_for_batching._patched_call(inputs: mx.array, inputs_embeds: Optional[mx.array] = None, cache: Optional[Any] = None, mask: Optional[mx.array] = None, position_ids: Optional[mx.array] = None) -> mx.array |
Nested Function patch_glm4v_moe_for_batching._patched_call calls self.embed_tokens, inputs_embeds.astype, isinstance, int; returns self.norm(h). |
#L50-L84 |