BlueSkills

Inspect an agent skill before you install it.

Review its instructions, bundled files and the evidence behind the verdict.

How BlueSkills analyzes an AI agent skill

What evidence should you review before allowing an AI agent to install a skill?

New Scan

BlueSkills helps a user answer one specific question: what evidence should I review before allowing an AI agent to install this skill?

It examines the submitted material, identifies suspicious instructions or capabilities, and returns the evidence behind its result. BlueSkills does not certify that a skill is safe.

What BlueSkills accepts

Pasted SKILL.md

A pasted SKILL.md is treated as a single-file submission, up to 1 MiB. Referenced scripts, hooks, manifests, remote content, and surrounding files are not included unless they are separately submitted.

Public GitHub or GitLab repository

BlueSkills retrieves the submitted public repository. The report records the commit the host names on the downloaded archive when the host names one. When a repository contains several skills, each skill is examined separately, up to five units on a hosted scan. Files outside those skill directories are also considered as package material. Further units are reported as unscanned, and an unscanned unit is not a clean result.

Private repositories are not fetched.

ZIP package

BlueSkills accepts an archive up to 25 MiB and extracts it within a 256 MiB unpacked limit, 10,000 files, and a per-file text cap of 4 MiB. A file over that text cap is truncated and the report says so.

An archive that contains a symlink, a special file, a duplicate or case-colliding member, or an encrypted archive is rejected. That rejection is an invalid submission, not a clean scan. Missing git submodules, oversized members, and truncated coverage are reported as gaps. A gap is not clean content.

How the analysis works

BlueSkills combines complementary forms of analysis. The report lists the stages that ran.

These stages run on every hosted scan. If one of them fails, the scan cannot be a complete CLEAN result:

  • package and scope inspection, including files outside declared skill directories
  • text normalization and obfuscation handling
  • instruction and configuration analysis
  • bundled code and manifest analysis for the languages and manifests the scanner supports
  • dependency and install-time checks

These stages are enabled on the hosted worker and are conditional:

  • advisory lookup for declared dependencies, when the lookup succeeds
  • an additional signature and data-flow pass shipped with the worker
  • model-assisted review of instructions and of similarity to known-malicious snippets
  • a final model review, only while the result is still SUSPICIOUS and the scan has no critical finding; that review can raise SUSPICIOUS to MALICIOUS and cannot clear a finding
  • isolated runtime observation, only when the deployment has that stage turned on, and skipped when the earlier result is already MALICIOUS

The worker image leaves isolated runtime observation off unless the running job turns it on. When the stage is on and it fails, times out, or has no capacity, CLEAN is not a possible complete result.

An embedding pass that was rate-limited or unavailable also cannot be a complete CLEAN result. A failed optional lookup that is only recorded as a note does not, by itself, change a CLEAN result.

Package and scope inspection

The scanner determines which files belong to each skill and which files exist outside declared skill directories. This prevents a clean SKILL.md from hiding a harmful installer, hook, or adjacent script.

Text normalization and obfuscation handling

BlueSkills normalizes text and looks for techniques that conceal instructions, including encoded content, invisible characters, and suspicious script or language changes.

Instruction and configuration analysis

The scanner examines skill instructions, bundled scripts, dependency manifests, and other material capable of changing how an agent behaves.

Bundled code and manifest analysis

BlueSkills examines supported bundled scripts and dependency manifests for dangerous execution capabilities, credential access, unexpected outbound communication, install-time behavior, unsafe execution, and supply-chain warning signs.

Model-assisted analysis

Selected analysis layers use machine-learning or language-model review to identify related or contextual behavior that deterministic checks may not fully express.

Model-assisted output is evidence with limitations. It does not turn incomplete coverage into proof of safety.

Isolated runtime observation

When this stage is enabled, and the earlier result is not already malicious, BlueSkills can allow an agent to interact with the submitted skill inside an isolated environment containing planted credentials and monitored network access.

Runtime findings are observations from one environment and one execution. They do not prove how every operating system, user state, timing condition, or future execution path will behave. An uneventful run is not proof of safety.

Understanding the report

CLEAN

BlueSkills found no sufficiently suspicious behavior in the material and analysis it covered. CLEAN is not a certificate that the skill is safe.

SUSPICIOUS

The scan found warning signs requiring manual review. Read the quoted evidence, inspect the requested capabilities, and consider the access the skill would receive before deciding whether to install it.

MALICIOUS

The evidence indicates harmful behavior. BlueSkills recommends against installing or running the skill.

Risk score

The risk score helps order and prioritize reports. It is not a percentage probability and must not be interpreted as a percentage-safe measurement.

Coverage

The report states what was examined, which analysis layers ran, and where coverage was incomplete. An incomplete analysis must not be treated as a clean result. Required-stage failure, a coverage hole, an unavailable embedding pass, and a failed runtime stage when runtime is enabled all block a complete CLEAN result.

Known limitations

BlueSkills cannot establish that:

  • a skill will behave identically on every operating system
  • every execution path was reached during runtime observation
  • remote content fetched after installation is safe
  • the submitted source will remain unchanged
  • future versions will behave like the scanned revision
  • encrypted or previously unseen behavior will always be recognized
  • a CLEAN result makes installation risk-free
  • planted credential values will be recognized after they have been compressed or otherwise transformed before they leave the isolated environment

Scan the exact revision you intend to install and review the skill's requested access independently.

Public adversarial testing

BlueSkills has been tested with malicious examples, benign controls, internal regression tests, and public challenge submissions.

The first public challenge produced three prize-winning misses. The status below reflects the source tree after the 2026-09-24 fixes.

  1. Documentation that rewires deployment addresses. Partially mitigated. An Ethereum address in the skill is now reported as a warning, so this submission no longer receives a complete CLEAN result. The warning does not establish that a particular configuration edit is malicious. Regression: test_warns_on_ethereum_address. Maintainer rescan recorded on the issue: SUSPICIOUS.
  2. macOS credential exfiltration that stayed quiet on Linux. Patched in source. Code that builds a credential path and also performs network access is reported, and a skill that targets macOS is observed with the runtime reporting Darwin. Regressions: test_reconstructed_ssh_key_with_network_is_ast013, test_str_join_ssh_key_split_across_files_is_ast013, test_needs_darwin_for_compatibility_or_platform_branch, test_darwin_sitecustomize_reports_darwin.
  3. Generic dotfile collection packed into an archive and uploaded. Partially mitigated. The runtime agent now follows the skill even when the description tells it to wait for a user phrase. Regression: test_agent_prompt_runs_the_skill_and_every_branch. Recognition of planted credential values inside a compressed upload is still an open limitation. This issue is not fixed.

Why the exact analyzers are not public

BlueSkills publishes its purpose, high-level architecture, result semantics, limitations, and lessons from adversarial testing. Its exact detection rules, thresholds, and normalization details remain private.

Publishing the complete detection checklist would allow an attacker to optimize malicious material against known checks. Keeping those details private does not make BlueSkills unbreakable, which is why reproducible failures and public challenge submissions remain part of the evaluation process.

Current version and evidence

  • CLI: bluethroat on PyPI (uv tool install bluethroat).
  • Methodology last updated: 2026-09-28
  • Supported inputs: pasted SKILL.md (1 MiB), public GitHub URL, public GitLab URL, ZIP (25 MiB)
  • Maximum submission limits: 25 MiB archive, 256 MiB unpacked, 10,000 files, 4 MiB of text per file, 5 skill units on a hosted scan
  • Runtime environment: isolated sandbox with planted credentials and monitored outbound traffic, only when the worker's runtime stage is enabled. The image default leaves that stage off. The live job must be checked before claiming that runtime runs in production.
  • Public architecture overview: github.com/BluethroatLabs/blueskills-public

Challenge a result

If you find a reproducible missed detection or incorrect finding, submit it through the BlueSkills public issue tracker.

Use dummy data. Include the exact source, revision or ZIP, scan ID, observed result, expected result, and a negative control when possible. Do not publish real credentials or private packages.