Skip to main content

New: Malicious MCP server caught in the wild!

Stop malicious prompts at the source

Instructions hidden in files, skills, web pages, and MCP responses hijack agents into leaking data and running harmful commands. Kirin catches the injection and blocks the action it would have caused.

The injection is the input. The damage is the action.

An agent cannot tell a README from a command. Telling the model to ignore malicious instructions does not work; it agrees, then follows them. Kirin inspects what the agent is about to do, not just what it was told, and stops the dangerous action regardless of where the instruction came from.

How Kirin defends against prompt injection

Hidden instructions in skills

Before Knostic

A skill's SKILL.md carries a hidden instruction. The agent follows it, ships ~/.aws/credentials to a remote server, and reports success.

After Knostic

AgentMesh flags the skill. Kirin blocks the install and, if a hidden instruction slips through elsewhere, the exfiltrating command.

Injected destructive commands

Before Knostic

A disguised instruction tells the agent to drop a production table. It does.

After Knostic

The command is evaluated against policy and denied before execution, with the alert logged as prevented.

Key capabilities

Real-time threat scanning

Prompts, retrieved context, and tool responses are scanned for injection and malicious instructions.

Action-level blocking

Dangerous commands and data movements are stopped before they execute, whatever their origin.

Secret detection

Credentials in prompts and outputs caught and blocked or redacted, in Monitor or Enforce.

Supply-chain scanning

AgentMesh detects prompt injection in skills, MCP servers, and extensions before install.

Tunable policies

Choose block, downgrade, route to a human, or log per rule and per team.

Audit logs

Every detection and decision recorded for incident response.

Frequently asked questions

Content crafted to make an AI agent override its instructions, exfiltrate data, or take harmful actions. It arrives through anything the agent reads: files, web pages, tool outputs, skills.

Filters and firewalls parse network traffic, not natural language inside an agent's context. The malicious instruction looks like text until the agent acts on it.

By scanning inputs for known injection patterns and, more importantly, by evaluating the resulting action against policy before it runs.

No. Policies are tuned per team, and only actions that violate them are interrupted.

Cursor, Claude Code, GitHub Copilot, Windsurf, Codex, JetBrains, Devin Desktop, Gemini CLI, and Claude Cowork.