Skip to content

vllm_mlx.tool_parsers.llama_tool_parser

Llama tool call parser for vllm-mlx.

View the complete module source at #L1-L128.

API details

Each callable below includes its exact signature, type annotations, inputs, defaults, return contract, documented exceptions, implementation source, and parsed docstring sections when the source provides them.

vllm_mlx.tool_parsers.llama_tool_parser

Llama tool call parser for vllm-mlx.

Handles Llama's tool calling format: - XML style: {"arg": "value"}

vllm_mlx.tool_parsers.llama_tool_parser.LlamaToolParser

LlamaToolParser(tokenizer: PreTrainedTokenizerBase | None = None)

Bases: ToolParser

Tool call parser for Llama models.

Supports Llama tool call format: - {"arg": "value"}

Used when --enable-auto-tool-choice --tool-call-parser llama are set.

Source code in vllm_mlx/tool_parsers/abstract_tool_parser.py
def __init__(self, tokenizer: PreTrainedTokenizerBase | None = None):
    """
    Initialize the tool parser.

    Args:
        tokenizer: The tokenizer for the model (optional, some parsers need it)
    """
    self.model_tokenizer = tokenizer
    # State for streaming parsing
    self.current_tool_id: int = -1
    self.prev_tool_call_arr: list[dict] = []

vllm_mlx.tool_parsers.llama_tool_parser.LlamaToolParser.SUPPORTS_NATIVE_TOOL_FORMAT class-attribute instance-attribute

SUPPORTS_NATIVE_TOOL_FORMAT = True

vllm_mlx.tool_parsers.llama_tool_parser.LlamaToolParser.FUNCTION_PATTERN class-attribute instance-attribute

FUNCTION_PATTERN = re.compile('<function=([^>]+)>(\\{.*?\\})</function>', re.DOTALL)

vllm_mlx.tool_parsers.llama_tool_parser.LlamaToolParser.extract_tool_calls

extract_tool_calls(model_output: str, request: dict[str, Any] | None = None) -> ExtractedToolCallInformation

Extract tool calls from a complete Llama model response.

Source code in vllm_mlx/tool_parsers/llama_tool_parser.py
def extract_tool_calls(
    self, model_output: str, request: dict[str, Any] | None = None
) -> ExtractedToolCallInformation:
    """
    Extract tool calls from a complete Llama model response.
    """
    tool_calls = []
    cleaned_text = model_output

    matches = self.FUNCTION_PATTERN.findall(model_output)
    for name, args_str in matches:
        try:
            arguments = json.loads(args_str)
            tool_calls.append(
                {
                    "id": generate_tool_id(),
                    "name": name.strip(),
                    "arguments": (
                        json.dumps(arguments, ensure_ascii=False)
                        if isinstance(arguments, dict)
                        else str(arguments)
                    ),
                }
            )
        except json.JSONDecodeError:
            # Keep the raw arguments string
            tool_calls.append(
                {
                    "id": generate_tool_id(),
                    "name": name.strip(),
                    "arguments": args_str,
                }
            )

    if matches:
        cleaned_text = self.FUNCTION_PATTERN.sub("", cleaned_text).strip()

    if tool_calls:
        return ExtractedToolCallInformation(
            tools_called=True,
            tool_calls=tool_calls,
            content=cleaned_text if cleaned_text else None,
        )
    else:
        return ExtractedToolCallInformation(
            tools_called=False, tool_calls=[], content=model_output
        )

vllm_mlx.tool_parsers.llama_tool_parser.LlamaToolParser.extract_tool_calls_streaming

extract_tool_calls_streaming(previous_text: str, current_text: str, delta_text: str, previous_token_ids: Sequence[int] | None = None, current_token_ids: Sequence[int] | None = None, delta_token_ids: Sequence[int] | None = None, request: dict[str, Any] | None = None) -> dict[str, Any] | None

Extract tool calls from streaming Llama model output.

Source code in vllm_mlx/tool_parsers/llama_tool_parser.py
def extract_tool_calls_streaming(
    self,
    previous_text: str,
    current_text: str,
    delta_text: str,
    previous_token_ids: Sequence[int] | None = None,
    current_token_ids: Sequence[int] | None = None,
    delta_token_ids: Sequence[int] | None = None,
    request: dict[str, Any] | None = None,
) -> dict[str, Any] | None:
    """
    Extract tool calls from streaming Llama model output.
    """
    # Check for tool call markers
    if "<function=" not in current_text:
        return {"content": delta_text}

    # If we detect end of function, parse
    if "</function>" in delta_text:
        result = self.extract_tool_calls(current_text)
        if result.tools_called:
            return {
                "tool_calls": [
                    {
                        "index": i,
                        "id": tc["id"],
                        "type": "function",
                        "function": {
                            "name": tc["name"],
                            "arguments": tc["arguments"],
                        },
                    }
                    for i, tc in enumerate(result.tool_calls)
                ]
            }

    return None

vllm_mlx.tool_parsers.llama_tool_parser.generate_tool_id

generate_tool_id() -> str

Generate a unique tool call ID.

Source code in vllm_mlx/tool_parsers/llama_tool_parser.py
def generate_tool_id() -> str:
    """Generate a unique tool call ID."""
    return f"call_{uuid.uuid4().hex[:8]}"

Complete contract reference

Expand any definition for its exact inputs, annotations, defaults, return contract, directly raised exceptions, source-grounded behavior, and immutable line link. This section includes private and nested definitions that ordinary API generators omit.

vllm_mlx.tool_parsers.llama_tool_parser.generate_tool_id · function
vllm_mlx.tool_parsers.llama_tool_parser.generate_tool_id() -> str

Generate a unique tool call ID.

Parameters

This callable has no explicit inputs.

Returns

  • Type: str
  • Direct return expressions: f'call_{uuid.uuid4().hex[:8]}'

Exceptions and behavior

Function generate_tool_id calls uuid.uuid4; returns f'call_{uuid.uuid4().hex[:8]}'. No direct raise statement appears in this definition.

View source #L22-L24.

vllm_mlx.tool_parsers.llama_tool_parser.LlamaToolParser · class
vllm_mlx.tool_parsers.llama_tool_parser.LlamaToolParser()

Tool call parser for Llama models.

Parameters

This callable has no explicit inputs.

Returns

  • Constructs: vllm_mlx.tool_parsers.llama_tool_parser.LlamaToolParser

Exceptions and behavior

Class LlamaToolParser derives from ToolParser and declares 2 direct member(s). No direct raise statement appears in this definition.

View source #L28-L128.

vllm_mlx.tool_parsers.llama_tool_parser.LlamaToolParser.extract_tool_calls · method
vllm_mlx.tool_parsers.llama_tool_parser.LlamaToolParser.extract_tool_calls(model_output: str, request: dict[str, Any] | None = None) -> ExtractedToolCallInformation

Extract tool calls from a complete Llama model response.

Parameters

Name Type Required Default Description
model_output str yes none Required positional or keyword input.
request dict[str, Any] \| None no None Optional positional or keyword input; defaults to None.

Returns

  • Type: ExtractedToolCallInformation
  • Direct return expressions: ExtractedToolCallInformation(tools_called=True, tool_calls=tool_calls, content=cleaned_text if cleaned_text else None); ExtractedToolCallInformation(tools_called=False, tool_calls=[], content=model_output)

Exceptions and behavior

Method LlamaToolParser.extract_tool_calls calls self.FUNCTION_PATTERN.findall, json.loads, tool_calls.append, generate_tool_id; has 2 explicit return paths. No direct raise statement appears in this definition.

View source #L44-L90.

vllm_mlx.tool_parsers.llama_tool_parser.LlamaToolParser.extract_tool_calls_streaming · method
vllm_mlx.tool_parsers.llama_tool_parser.LlamaToolParser.extract_tool_calls_streaming(previous_text: str, current_text: str, delta_text: str, previous_token_ids: Sequence[int] | None = None, current_token_ids: Sequence[int] | None = None, delta_token_ids: Sequence[int] | None = None, request: dict[str, Any] | None = None) -> dict[str, Any] | None

Extract tool calls from streaming Llama model output.

Parameters

Name Type Required Default Description
previous_text str yes none Required positional or keyword input.
current_text str yes none Required positional or keyword input.
delta_text str yes none Required positional or keyword input.
previous_token_ids Sequence[int] \| None no None Optional positional or keyword input; defaults to None.
current_token_ids Sequence[int] \| None no None Optional positional or keyword input; defaults to None.
delta_token_ids Sequence[int] \| None no None Optional positional or keyword input; defaults to None.
request dict[str, Any] \| None no None Optional positional or keyword input; defaults to None.

Returns

  • Type: dict[str, Any] | None
  • Direct return expressions: {'content': delta_text}; {'tool_calls': [{'index': i, 'id': tc['id'], 'type': 'function', 'function': {'name': tc['name'], 'arguments': tc['argu…; None

Exceptions and behavior

Method LlamaToolParser.extract_tool_calls_streaming calls self.extract_tool_calls, enumerate; has 3 explicit return paths. No direct raise statement appears in this definition.

View source #L92-L128.

Complete symbol map

This map also includes private definitions and nested helpers. The signature column exposes every explicit input even when an internal helper has no dedicated parameter prose.

Symbol Kind Signature and inputs What it does Source
generate_tool_id function generate_tool_id() -> str Generate a unique tool call ID. #L22-L24
LlamaToolParser class LlamaToolParser() Tool call parser for Llama models. #L28-L128
LlamaToolParser.extract_tool_calls method LlamaToolParser.extract_tool_calls(model_output: str, request: dict[str, Any] \| None = None) -> ExtractedToolCallInformation Extract tool calls from a complete Llama model response. #L44-L90
LlamaToolParser.extract_tool_calls_streaming method LlamaToolParser.extract_tool_calls_streaming(previous_text: str, current_text: str, delta_text: str, previous_token_ids: Sequence[int] \| None = None, current_token_ids: Sequence[int] \| None = None, delta_token_ids: Sequence[int] \| None = None, request: dict[str, Any] \| None = None) -> dict[str, Any] \| None Extract tool calls from streaming Llama model output. #L92-L128