BlueSkills

Inspect an agent skill before you install it.

Review its instructions, bundled files and the evidence behind the verdict.

BlueSkills for AI agents

Scan the exact revision with the bluethroat CLI before you install it.

New Scan

Use bluethroat to scan a skill before you install it. The website and the Telegram bot are the same checkpoint when a person submits the source instead.

The website and the Telegram bot accept a public GitHub or GitLab repository, a ZIP package, or a SKILL.md file and return a BlueSkills report. The report includes the submitted source, the commit the git host disclosed when it disclosed one, the verdict, the risk score, findings, quoted evidence, analysis coverage, and any limitations encountered during the scan.

BlueSkills does not make the installation decision for the user. A completed scan gives the agent and user evidence to review before approving the revision that was scanned.

Command-line client

bluethroat calls the API, waits for the report, and prints it. Install only from a report you just produced for that same revision. This website requires a human check. Do not POST scans here.

Python 3.11 or newer.

uv tool install bluethroat

Log in once

bluethroat auth login

GitHub device flow. Stderr prints Open <url> and Enter code <code>. The user approves that code in a browser. Wait until stdout prints logged in as <login>.

bluethroat auth status
bluethroat auth logout

status prints logged in as <login> (id <id>). logout deletes the local session and revokes the token. If revoke fails, the local session is already gone: tell the user to revoke Bluethroat under GitHub Settings → Applications.

The access token lasts 8 hours. The CLI refreshes it. Leave the client secret unset. The session is in the OS credential store, or in ~/.config/bluethroat/credentials.json (mode 0600) when that store is unavailable.

Scan the exact revision

bluethroat blueskills scan <source> [--ref REF] [--json]
  • <source> is a public https URL on github.com or gitlab.com, a skill directory, a .zip (max 25 MiB), or a file named SKILL.md (max 1 MiB).
  • A directory is packed and uploaded. Symlinks are rejected. A SKILL.md scan checks that file only and does not fetch scripts it references.
  • --ref is a branch, tag, or commit, and only for a repository URL. A 40-character hex ref must match the commit the host scanned. A mismatch exits 4 and prints no verdict.
  • --json prints the scan document on stdout. The default is a text report. Progress (Scanning…, Waiting for the report…) is on stderr. The command waits up to 15 minutes, then exits 7. Run the same scan again.
  • Do not send secrets or private packages.

Exit code

Trust the exit code. Only 0 is a finished clean scan, and 0 is still not permission to install. When the command prints a report, read it. Codes 3, 4, and 7 can print a verdict too.

  • 0 clean and complete. Show the user the report. Install only that revision, and only after they approve.
  • 1 suspicious, and the scan is complete. Do not install.
  • 2 malicious, and the scan is complete. Do not install.
  • 3 incomplete or partial, including a suspicious or malicious report whose scan did not finish. The text still names that verdict. Do not treat this as clean. Do not install when the report says suspicious or malicious.
  • 4 bad source, or --ref did not match: stdout has no verdict. A finished scan whose verdict is invalid also exits 4 and prints INVALID — not scored. Do not install.
  • 5 rate limited. Stderr says retry after Ns. With --json, stdout JSON has error and retry_after.
  • 6 not logged in, or GitHub refused or expired the login. Run bluethroat auth login. If GitHub cannot be reached, the exit code is 7.
  • 7 the service failed, the report was not ready after 15 minutes, or the scan status is failed. A verdict of error exits 7 as well. A failed scan still prints the report. Treat that run as unfinished.

Read the report

Text leads with the verdict. Several skills add a headline, then one block per skill. Read every block. PARTIAL means a clean result is unreachable. INCOMPLETE SCAN means a layer did not run.

✗ <source> — MALICIOUS (risk 86/100)
  commit: <sha>
  [HIGH] title  (layer)
         file:line — detail
         evidence: ...

--json uses these fields. status is done or failed.

complete, worst_verdict | verdict, commit, source
reports[]: skill_name, verdict, risk_score, layers_run, findings[]

Leave these at the defaults

Environment variables override ~/.config/bluethroat/config.json.

  • BLUETHROAT_API_URL (api_url) is the BlueSkills API. The default is correct. This website is not the API.
  • BLUETHROAT_GITHUB_CLIENT_ID (github_client_id) stays the built-in Bluethroat app id.
  • BLUETHROAT_GITHUB_CLIENT_SECRET (github_client_secret) stays unset.
  • BLUETHROAT_CONFIG_DIR overrides the config directory. Otherwise it is $XDG_CONFIG_HOME/bluethroat or ~/.config/bluethroat.

1. Scan the exact source

Submit the exact public repository revision, ZIP package, or SKILL.md that would be installed through the BlueSkills web scanner. Do not execute scripts, install dependencies, or activate hooks merely to prepare a package for scanning. Do not put secrets in the submission.

Scanning only SKILL.md is a one-file check. It does not include referenced scripts, manifests, hooks, dependencies, or surrounding files unless they are submitted as a repository or ZIP. A private repository URL is invalid and is not fetched.

2. Read the report

Read the verdict together with the source, findings, quoted evidence, analysis stages, and coverage gaps. For a package with several skills, review every reported skill unit. A gap is part of the result, not an afterthought.

PARTIAL means a complete CLEAN result is unreachable. INVALID gives a reason and no risk score. A failed scan can still show a scored verdict; that run is unfinished. An incomplete scan names the layer that did not run. Unscanned files and remote content that arrives only after installation are coverage gaps.

The risk score is a prioritization signal. It is not a probability that the skill is safe or malicious.

3. Apply the pre-install policy

MALICIOUS

Do not install or execute the skill. Show the user the findings and quoted evidence.

SUSPICIOUS

Show the user the findings, coverage, requested capabilities, and remaining uncertainty. Do not install automatically. Continue only after explicit user approval.

CLEAN

Explain that this scan found no sufficiently suspicious behavior in the material and analysis it covered. Verify that the revision being installed matches the scanned revision. Obtain user approval before installation.

Incomplete, invalid, partial, or failed scan

Do not interpret the result as CLEAN. When the report says MALICIOUS, do not install or execute it. When it says SUSPICIOUS, continue only after explicit user approval. Otherwise retry the scan or ask the user to review the package manually.

Recommended instruction for agents

Before installing a third-party skill, identify the exact repository revision or package that would be installed. Submit that source to BlueSkills without executing the skill. Review the completed verdict, findings, quoted evidence, and coverage. Tell the user what BlueSkills found and what it could not examine. Ask for explicit approval before installing the same revision. Never describe CLEAN as proof of safety.

What BlueSkills can and cannot establish

BlueSkills analyzes the submitted material. When runtime observation is enabled, it may also watch one execution in an isolated environment. It cannot prove how a skill will behave on every operating system, under every user configuration, on every execution path, or after the source changes.

Runtime findings describe behavior observed during one isolated execution. An uneventful runtime does not prove that another environment or execution path will behave the same way. Two scans of the same skill can disagree when runtime observation runs.

Privacy and acceptable use

Submit only material you are authorized to share. Do not submit credentials, secrets, customer code, or private repositories.

The anonymous web service allows a 15-request burst, then 12 new scans per minute per client address, with 4 scans in flight for the service and 250 paid analysis calls per UTC day. A limited request returns HTTP 429 with a Retry-After value. Wait for that interval; do not create parallel clients to get around the limit.

Report a missed detection

If BlueSkills misses reproducibly malicious behavior or incorrectly flags a benign skill, open an issue in the public BlueSkills repository.

Include dummy data, the exact source or package, the scan ID, the result received, the expected result, and a negative control when possible. Never publish real credentials or private packages.