# `vllm_mlx.patches.gemma4_mllm`

Runtime patch for mlx-vlm Gemma 4 Attention to trim oversized masks.

[View the complete module source at #L1-L98](https://github.com/waybarrios/vllm-mlx/blob/a69d47912bcb21d8fe04d48f75fa896b620ffcfa/vllm_mlx/patches/gemma4_mllm.py#L1-L98).

## API details

Each callable below includes its exact signature, type annotations, inputs, defaults, return contract, documented exceptions, implementation source, and parsed docstring sections when the source provides them.

::: vllm_mlx.patches.gemma4_mllm
    options:
      members:
        - logger
        - patch_gemma4_attention_for_batching
      filters: []
      show_if_no_docstring: true

## Complete contract reference

Expand any definition for its exact inputs, annotations, defaults, return contract, directly raised exceptions, source-grounded behavior, and immutable line link. This section includes private and nested definitions that ordinary API generators omit.

<details class="api-contract" id="contract-vllm_mlx.patches.gemma4_mllm.patch_gemma4_attention_for_batching" markdown="1">
<summary><code>vllm_mlx.patches.gemma4_mllm.patch_gemma4_attention_for_batching</code> · function</summary>

```python
vllm_mlx.patches.gemma4_mllm.patch_gemma4_attention_for_batching() -> bool
```

Patch Gemma 4 Attention.__call__ to trim oversized masks.

**Parameters**

This callable has no explicit inputs.

**Returns**

- Type: `bool`
- Direct return expressions: `False`; `True`

**Exceptions and behavior**

Function `patch_gemma4_attention_for_batching` calls `logger.debug`, `getattr`, `logger.info`; has 2 explicit return paths.
No direct `raise` statement appears in this definition.

[View source #L28-L98](https://github.com/waybarrios/vllm-mlx/blob/a69d47912bcb21d8fe04d48f75fa896b620ffcfa/vllm_mlx/patches/gemma4_mllm.py#L28-L98).

</details>

<details class="api-contract" id="contract-vllm_mlx.patches.gemma4_mllm.patch_gemma4_attention_for_batching._patched_call" markdown="1">
<summary><code>vllm_mlx.patches.gemma4_mllm.patch_gemma4_attention_for_batching._patched_call</code> · nested function</summary>

```python
vllm_mlx.patches.gemma4_mllm.patch_gemma4_attention_for_batching._patched_call(x: mx.array, mask: Optional[mx.array] = None, cache: Optional[Any] = None, shared_kv: Optional[tuple] = None, offset: Optional[Any] = None) -> Any
```

Nested Function `patch_gemma4_attention_for_batching._patched_call` calls `self.q_proj(x).reshape`, `self.q_proj`, `self.q_norm`, `self.k_proj(x).reshape`; returns `(self.o_proj(output), (keys, values), offset)`.

**Parameters**

| Name | Type | Required | Default | Description |
| --- | --- | --- | --- | --- |
| `x` | `mx.array` | `yes` | `none` | Required positional or keyword input. |
| `mask` | `Optional[mx.array]` | `no` | `None` | Optional positional or keyword input; defaults to `None`. |
| `cache` | `Optional[Any]` | `no` | `None` | Optional positional or keyword input; defaults to `None`. |
| `shared_kv` | `Optional[tuple]` | `no` | `None` | Optional positional or keyword input; defaults to `None`. |
| `offset` | `Optional[Any]` | `no` | `None` | Optional positional or keyword input; defaults to `None`. |

**Returns**

- Type: `Any`
- Direct return expressions: `(self.o_proj(output), (keys, values), offset)`

**Exceptions and behavior**

Nested Function `patch_gemma4_attention_for_batching._patched_call` calls `self.q_proj(x).reshape`, `self.q_proj`, `self.q_norm`, `self.k_proj(x).reshape`; returns `(self.o_proj(output), (keys, values), offset)`.
No direct `raise` statement appears in this definition.

[View source #L45-L93](https://github.com/waybarrios/vllm-mlx/blob/a69d47912bcb21d8fe04d48f75fa896b620ffcfa/vllm_mlx/patches/gemma4_mllm.py#L45-L93).

</details>

## Complete symbol map

This map also includes private definitions and nested helpers. The signature column exposes every explicit input even when an internal helper has no dedicated parameter prose.

| Symbol | Kind | Signature and inputs | What it does | Source |
| --- | --- | --- | --- | --- |
| [`patch_gemma4_attention_for_batching`](#contract-vllm_mlx.patches.gemma4_mllm.patch_gemma4_attention_for_batching) | function | `patch_gemma4_attention_for_batching() -> bool` | Patch Gemma 4 Attention.__call__ to trim oversized masks. | [#L28-L98](https://github.com/waybarrios/vllm-mlx/blob/a69d47912bcb21d8fe04d48f75fa896b620ffcfa/vllm_mlx/patches/gemma4_mllm.py#L28-L98) |
| [`patch_gemma4_attention_for_batching._patched_call`](#contract-vllm_mlx.patches.gemma4_mllm.patch_gemma4_attention_for_batching._patched_call) | nested function | `patch_gemma4_attention_for_batching._patched_call(x: mx.array, mask: Optional[mx.array] = None, cache: Optional[Any] = None, shared_kv: Optional[tuple] = None, offset: Optional[Any] = None) -> Any` | Nested Function `patch_gemma4_attention_for_batching._patched_call` calls `self.q_proj(x).reshape`, `self.q_proj`, `self.q_norm`, `self.k_proj(x).reshape`; returns `(self.o_proj(output), (keys, values), offset)`. | [#L45-L93](https://github.com/waybarrios/vllm-mlx/blob/a69d47912bcb21d8fe04d48f75fa896b620ffcfa/vllm_mlx/patches/gemma4_mllm.py#L45-L93) |
