llm-semantic-router / toolcall-verifier

huggingface.co
Total runs: 44
24-hour runs: 0
7-day runs: 28
30-day runs: 28
Model's Last Updated: December 18 2025
token-classification

Introduction of toolcall-verifier

Model Details of toolcall-verifier

ToolCallVerifier - Unauthorized Tool Call Detection

License Model

Stage 2 of Two-Stage LLM Agent Defense Pipeline


๐ŸŽฏ What This Model Does

ToolCallVerifier is a ModernBERT-based token classifier that detects unauthorized tool calls in LLM agent systems. It performs token-level classification on tool call JSON to identify malicious arguments that may have been injected through prompt injection attacks.

Label Description
AUTHORIZED Token is part of a legitimate, user-requested action
UNAUTHORIZED Token indicates injected/malicious content โ€” BLOCK

๐Ÿšจ Attack Categories Covered
Category Source Description
Delimiter Injection LLMail <<end_context>> , >>}}\]\])
Word Obfuscation LLMail Inserting noise words between tokens
Fake Sessions LLMail START_USER_SESSION , EXECUTE_USERQUERY
Roleplay Injection WildJailbreak "You are an admin bot that can..."
XML Tag Injection WildJailbreak <execute_action> , <tool_call>
Authority Bypass WildJailbreak "As administrator, I authorize..."
Intent Mismatch Synthetic User asks X, tool does Y
MCP Tool Poisoning Synthetic Hidden exfiltration in tool args
MCP Shadowing Synthetic Fake authorization context
๐Ÿ”— Integration with FunctionCallSentinel

This model is Stage 2 of a two-stage defense pipeline:

โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”     โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”     โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
โ”‚   User Prompt   โ”‚โ”€โ”€โ”€โ”€โ–ถโ”‚ ToolCallSentinel โ”‚โ”€โ”€โ”€โ”€โ–ถโ”‚   LLM + Tools   โ”‚
โ”‚                 โ”‚     โ”‚      (Stage 1)       โ”‚     โ”‚                 โ”‚
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜     โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜     โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
                                                              โ”‚
                               โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ–ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
                               โ”‚           ToolCallVerifier (This Model)                 โ”‚
                               โ”‚   Token-level verification before tool execution        โ”‚
                               โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
Scenario Recommendation
General chatbot Stage 1 only
Tool-calling agent (low risk) Stage 1 only
Tool-calling agent (high risk) Both stages
Email/file system access Both stages
Financial transactions Both stages

๐ŸŽฏ Intended Use
Primary Use Cases
  • LLM Agent Security : Verify tool calls before execution
  • Prompt Injection Defense : Detect unauthorized actions from injected prompts
  • API Gateway Protection : Filter malicious tool calls at infrastructure level
Out of Scope
  • General text classification
  • Non-tool-calling scenarios
  • Languages other than English
๐Ÿ“œ License

Apache 2.0

Runs of llm-semantic-router toolcall-verifier on huggingface.co

44
Total runs
0
24-hour runs
0
3-day runs
28
7-day runs
28
30-day runs

More Information About toolcall-verifier huggingface.co Model

More toolcall-verifier license Visit here:

https://choosealicense.com/licenses/apache-2.0

toolcall-verifier huggingface.co

toolcall-verifier huggingface.co is an AI model on huggingface.co that provides toolcall-verifier's model effect (), which can be used instantly with this llm-semantic-router toolcall-verifier model. huggingface.co supports a free trial of the toolcall-verifier model, and also provides paid use of the toolcall-verifier. Support call toolcall-verifier model through api, including Node.js, Python, http.

llm-semantic-router toolcall-verifier online free

toolcall-verifier huggingface.co is an online trial and call api platform, which integrates toolcall-verifier's modeling effects, including api services, and provides a free online trial of toolcall-verifier, you can try toolcall-verifier online for free by clicking the link below.

llm-semantic-router toolcall-verifier online free url in huggingface.co:

https://huggingface.co/llm-semantic-router/toolcall-verifier

toolcall-verifier install

toolcall-verifier is an open source model from GitHub that offers a free installation service, and any user can find toolcall-verifier on GitHub to install. At the same time, huggingface.co provides the effect of toolcall-verifier install, users can directly use toolcall-verifier installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

toolcall-verifier install url in huggingface.co:

https://huggingface.co/llm-semantic-router/toolcall-verifier

Url of toolcall-verifier

Provider of toolcall-verifier huggingface.co

llm-semantic-router
ORGANIZATIONS

Other API from llm-semantic-router