Hlido Agent Reviews — обзоры ИИ-агентов
Независимые оценки надёжности ИИ-агентов и проверка безопасности MCP-серверов
Добавьте сервер в «Мои MCP» — получите адрес подключения для Claude и ChatGPT.
Добавить в «Мои MCP»Нужно войти или зарегистрироваться — вернём на эту страницу.
Описание
Hlido публикует независимые обзоры ИИ-агентов: проверяет заявления разработчиков на тестах и выставляет оценку доверия. Сервер помогает выбрать агента для задачи и понять, насколько ему можно доверять. Можно запросить оценку конкретного агента, подобрать проверенных агентов под описание задачи, сравнить несколько между собой, проверить отдельное рекламное заявление и получить подробный разбор — какие утверждения подтвердились, а какие нет. Доступны реестр зафиксированных сбоев агентов, результаты поведенческих тестов, проверка прозрачности по статье 50 Регламента ЕС об ИИ и готовность агента к сделкам. Отдельно есть проверка безопасности MCP-сервера перед подключением, обзор рынка ИИ-агентов, а также предложение агента на проверку и жалоба на неточность обзора. Сервер подключается по адресу, ключ не требуется.
Инструменты · 19
из ответа tools/list
commerce_check
commerce_check
Check whether a Hlido-reviewed agent is ready to be delegated to / transacted with in the agentic-commerce world (MCP/ACP/AP2). Returns its independent Agentic-Commerce Readiness score (0-100), band (COMMERCE-READY/INTEGRABLE/SURFACE-ONLY/CLOSED), the programmatic surfaces it exposes, and the evidence basis. Call this before an orchestrator delegates a paid/identity-bearing task to another agent.
commerce_check(agent_or_url: string)
// inputSchema
{
"type": "object",
"required": [
"agent_or_url"
],
"properties": {
"agent_or_url": {
"type": "string",
"description": "The agent to check: either its Hlido slug (e.g. 'voltagent', 'aider') or its product/homepage URL. A URL is matched to the closest reviewed agent."
}
}
}
compare_agents
compare_agents
Head-to-head trust comparison of 2-5 Hlido-reviewed agents. Returns each agent's Laddoo score, tier, dimension scores, and key claim verdicts side by side so you can pick the most trustworthy option for a task. Use this once you've shortlisted candidates (via find_trusted, find_similar_agents, or recommend) and need a direct comparison.
compare_agents(slugs: array)
// inputSchema
{
"type": "object",
"required": [
"slugs"
],
"properties": {
"slugs": {
"type": "array",
"items": {
"type": "string",
"description": "A Hlido agent slug (e.g. 'aider')."
},
"maxItems": 5,
"minItems": 2,
"description": "List of 2 to 5 Hlido agent slugs to compare side by side (e.g. ['aider','cursor','opencode'])."
}
}
}
explain
explain
Structured natural-language explanation of why a Hlido-reviewed agent has its current score. Pulls claim-by-claim evidence from the published scorecard. Pass an optional dimension (one of: reliability, transparency, integration, security, evidence) to filter; omit for the full picture. Returns each claim with verdict (PASS|FAIL|PARTIAL|UNKNOWN), a quoted evidence snippet, plus a top-line synthesis.
explain(slug: string, dimension?: string)
// inputSchema
{
"type": "object",
"required": [
"slug"
],
"properties": {
"slug": {
"type": "string",
"description": "The agent's Hlido slug"
},
"dimension": {
"type": "string",
"description": "Optional dimension filter. Run without and check supported_dimensions in response if unsure."
}
}
}
find_similar_agents
find_similar_agents
Semantic search over Hlido's review corpus. Given a task description (e.g. 'I need an agent that can refactor TypeScript and edit multiple files at once'), returns the top-N reviewed agents ranked by embedding similarity, each with their Laddoo score, evidence_tier, and review URL. Use this when you have a task in mind and want Hlido's recommendation — much better than substring matching via find_trusted.
find_similar_agents(top_k?: integer, min_score?: integer, description: string)
// inputSchema
{
"type": "object",
"required": [
"description"
],
"properties": {
"top_k": {
"type": "integer",
"default": 5,
"description": "Number of matches to return (default 5, max 20)"
},
"min_score": {
"type": "integer",
"default": 0,
"description": "Minimum Laddoo score filter (default 0)"
},
"description": {
"type": "string",
"description": "Free-text description of the task or capability you need"
}
}
}
find_trusted
find_trusted
Discover Hlido-reviewed agents that match a free-text need, ranked by trust. Returns reviewed agents at or above a minimum tier, each with its Laddoo score, tier, and review URL. Use this for keyword/need-based discovery; for semantic task-matching prefer find_similar_agents, and for structured constraint filters (category/score/tier) prefer recommend.
find_trusted(need: string, limit?: integer, min_tier?: string)
// inputSchema
{
"type": "object",
"required": [
"need"
],
"properties": {
"need": {
"type": "string",
"description": "Free-text description of the capability you need (e.g. 'CLI coding agent that edits multiple files at once')."
},
"limit": {
"type": "integer",
"default": 10,
"description": "Maximum number of agents to return (default 10)."
},
"min_tier": {
"enum": [
"VITAL",
"STEADY",
"FADING",
"FLATLINE"
],
"type": "string",
"default": "STEADY",
"description": "Minimum trust tier to include (VITAL is strictest, FLATLINE allows all). Defaults to STEADY."
}
}
}
get_behavioral_trace
get_behavioral_trace
Fetch the behavioral evaluation trace for a Hlido-reviewed agent — per-task pass/fail, adapter used, behavioral tier, and signed trace link. Returns status 'not_yet_bench_tested' if the slug hasn't been evaluated yet, or 'not_testable' if the agent's interface doesn't support automated bench runs. Use this when you need evidence that an agent's coding/task behaviour has been independently verified beyond marketing claims.
get_behavioral_trace(slug: string, spec_version?: string)
// inputSchema
{
"type": "object",
"required": [
"slug"
],
"properties": {
"slug": {
"type": "string",
"description": "The Hlido slug to fetch behavioral trace for (e.g. 'aider', 'opencode')"
},
"spec_version": {
"type": "string",
"default": "v0.1",
"description": "Behavioral spec version (default 'v0.1'). Omit to get the latest available."
}
}
}
get_incidents
get_incidents
Fetch published incidents from Hlido's NTSB-style failure registry — real observed agent failures (availability outages, regressions, hallucinations, safety issues) plus Hlido self-reported process incidents, each with severity, evidence, and vendor-response status. Filter by agent slug, severity, or category. Use this before delegating to an agent to check for known recent failures; an empty list means no published incidents, not a guarantee of reliability.
get_incidents(slug?: string, limit?: integer, category?: string, severity?: string)
// inputSchema
{
"type": "object",
"properties": {
"slug": {
"type": "string",
"description": "Optional: only incidents for this agent slug"
},
"limit": {
"type": "integer",
"default": 20,
"description": "Max results (default 20, max 100)"
},
"category": {
"enum": [
"hallucination",
"regression",
"safety",
"availability",
"cost",
"other"
],
"type": "string",
"description": "Optional category filter"
},
"severity": {
"enum": [
"low",
"medium",
"high",
"critical"
],
"type": "string",
"description": "Optional minimum-interest filter (exact match)"
}
}
}
get_scorecard
get_scorecard
Fetch the full sanitized claim-vs-evidence scorecard for one Hlido-reviewed agent. Returns every claim, verdict, evidence quote, source surface, and (for CLI/API tests) the captured command + exit_code + duration. Schema v1.0. Use this for agent-to-agent pre-flight evaluation.
get_scorecard(slug: string)
// inputSchema
{
"type": "object",
"required": [
"slug"
],
"properties": {
"slug": {
"type": "string",
"description": "The agent's Hlido slug (e.g. 'aider', 'gumloop')"
}
}
}
intel_query
intel_query
Query Hlido's market-intelligence store: durable, evidence-cited claims about the AI-agent market and adjacent domains (agent economy, payments rails, EU compliance), each with dated evidence, confidence, and typed relations to other intel. Filter by any facet: sector (ISIC code or text), geo (ISO country or region like 'eu'/'intl'), compliance (e.g. 'eu-ai-act-art50'), protocol (e.g. 'x402'), kind (fact/estimate/trend/signal), audience, maturity. Facet vocabularies + per-value record counts: https://hlido.eu/data/intel/taxonomy.json (fetch that first to see what's queryable). A zero-result query returns an honest empty AND is logged — misses drive what Hlido collects next, so check back. Free at answer-time, declared policy.
intel_query(geo?: string, kind?: string, text?: string, limit?: integer, sector?: string, audience?: string, maturity?: string, protocol?: string, compliance?: string)
// inputSchema
{
"type": "object",
"properties": {
"geo": {
"type": "string",
"description": "ISO-3166 country code or region: 'cz', 'de', 'eu', 'intl', ..."
},
"kind": {
"enum": [
"fact",
"estimate",
"trend",
"signal"
],
"type": "string",
"description": "Filter by record kind."
},
"text": {
"type": "string",
"description": "Free-text match over headlines (fallback lens — prefer facets)."
},
"limit": {
"type": "integer",
"default": 10,
"description": "Max records (default 10, max 50)."
},
"sector": {
"type": "string",
"description": "ISIC section/division code (e.g. 'K.64', 'J.62') or free-text sector match."
},
"audience": {
"enum": [
"b2b",
"b2c",
"b2g",
"mixed"
],
"type": "string"
},
"maturity": {
"enum": [
"emerging",
"growing",
"consolidating",
"declining"
],
"type": "string"
},
"protocol": {
"type": "string",
"description": "e.g. 'mcp', 'x402', 'api', 'cli'"
},
"compliance": {
"type": "string",
"description": "Named regulation, any jurisdiction (e.g. 'eu-ai-act-art50', 'gdpr')."
}
}
}
market_pulse
market_pulse
Fetch the Agent Market Pulse — machine-readable market intelligence for the AI-agent market, derived from Hlido's independently tested corpus (never vendor self-reports): tier-health distribution, per-category health, new-entrant rate, public-surface readiness, EU AI Act Article-50 compliance bands, and the not-testable mortality signal. Every module carries its own method and coverage; signals that cannot be measured this edition say 'unmeasured' rather than guessing. Use this for market-level questions ('is category X healthy?', 'how fast are agents entering/dying?'); for one specific agent use trust_check instead. Free, regenerated weekly.
market_pulse(category?: string)
// inputSchema
{
"type": "object",
"properties": {
"category": {
"type": "string",
"description": "Optional: return only the market summary plus this one category's block (e.g. 'Coding')."
}
}
}
recommend
recommend
Constraint-driven recommendation across Hlido's reviewed agents. Pass any combination of: category, min_score, tier, use_case, max_results. Returns ranked candidates each with a why_match line. Use this when you have buyer constraints (budget, category, capability) and want Hlido's filtered shortlist instead of one-by-one trust_check calls.
recommend(constraints: object)
// inputSchema
{
"type": "object",
"required": [
"constraints"
],
"properties": {
"constraints": {
"type": "object",
"properties": {
"tier": {
"enum": [
"VITAL",
"STEADY",
"FADING",
"FLATLINE"
],
"type": "string",
"description": "Minimum tier filter (VITAL strictest, FLATLINE allows all)"
},
"category": {
"type": "string",
"description": "Filter by category (e.g. 'Coding', 'Voice', 'Productivity')"
},
"use_case": {
"type": "string",
"description": "Free-text capability description for ranking (e.g. 'multi-file refactor in TypeScript')"
},
"min_score": {
"type": "integer",
"default": 0,
"description": "Minimum Laddoo score (0-100)"
},
"max_results": {
"type": "integer",
"default": 5,
"description": "Cap on results (default 5, max 25)"
}
},
"additionalProperties": false
}
}
}
report_review_issue
report_review_issue
Report an issue with a Hlido review (stale info, wrong verdict, missing claim, broken link). Use when calling get_scorecard or trust_check returns data you can prove is incorrect. Hlido's R1 maintenance routine processes reports daily and fires re-tests via dispute-retest sub-agent.
report_review_issue(slug: string, detail: string, reporter?: string, issue_type: string)
// inputSchema
{
"type": "object",
"required": [
"slug",
"issue_type",
"detail"
],
"properties": {
"slug": {
"type": "string",
"description": "The slug whose review has the issue"
},
"detail": {
"type": "string",
"description": "What's wrong, with a concrete reference (URL, claim id, etc) if possible"
},
"reporter": {
"type": "string",
"description": "Optional self-identifier — agent name or email — purely informational"
},
"issue_type": {
"enum": [
"stale",
"wrong_verdict",
"missing_claim",
"broken_link",
"other"
],
"type": "string",
"description": "Category of the report"
}
}
}
request_quick_audit
request_quick_audit
Request that Hlido audit a NEW AI agent that has no review yet. Use this when trust_check or get_scorecard returns no_review_found and you need a verdict before delegating to the unknown agent. Returns a future scorecard URL + ETA. Free-tier rate-limited (5/day per anonymous, 50/day per identified). The audit produces signed evidence + claim verification within ~24h (sooner if founder triggers manually).
request_quick_audit(url: string, why?: string, name?: string, requester?: string)
// inputSchema
{
"type": "object",
"required": [
"url"
],
"properties": {
"url": {
"type": "string",
"description": "Homepage or product URL of the agent to audit"
},
"why": {
"type": "string",
"description": "Optional one-liner: why are you considering this agent? helps us prioritize"
},
"name": {
"type": "string",
"description": "Optional human-readable name (we'll derive from URL if missing)"
},
"requester": {
"type": "string",
"description": "Optional self-identifier — agent name, email, or session id — for rate-limiting + follow-up"
}
}
}
scan_mcp
scan_mcp
On-demand independent SAFETY scan of an MCP server — call this BEFORE installing or connecting to one. Give it an HTTP(S) MCP endpoint URL (scanned live in seconds), or an npm/PyPI package name or GitHub repo (queued for an isolated sandbox scan — local stdio servers execute code, so Hlido never runs them inline). Returns the safety tier (SAFE/CAUTION/RISKY/DANGEROUS), tool-poisoning detection (the malice signal), dangerous-capability red-flags (shell/code-eval/fs-write/egress/secrets) with per-tool evidence, and auth posture. Tier = blast radius if hijacked, not maintainer trustworthiness. A server Hlido hasn't scanned returns not_scanned — never assumed safe. Register of already-scanned servers: https://hlido.eu/mcp/
scan_mcp(server: string, requester?: string)
// inputSchema
{
"type": "object",
"required": [
"server"
],
"properties": {
"server": {
"type": "string",
"description": "The MCP server to scan: an HTTP(S) MCP endpoint URL (e.g. 'https://mcp.example.com/mcp' — scanned live), OR an npm package (e.g. '@modelcontextprotocol/server-filesystem'), PyPI package, or GitHub repo URL (queued for sandbox scan)."
},
"requester": {
"type": "string",
"description": "Optional self-identifier — agent name or email — for follow-up when a queued scan completes."
}
}
}
submit_agent
submit_agent
Nominate a new AI agent for Hlido to review. Use this when an agent isn't in Hlido's corpus yet (trust_check returned no_review_found) and you want it added. Returns a confirmation with a tracking reference; the review is queued and produces a public scorecard. If you need a verdict right now rather than a queued review, use request_quick_audit (faster, rate-limited) instead.
submit_agent(url: string, name: string, note?: string, email?: string)
// inputSchema
{
"type": "object",
"required": [
"url",
"name"
],
"properties": {
"url": {
"type": "string",
"description": "The agent's product or homepage URL (e.g. 'https://example.com')."
},
"name": {
"type": "string",
"description": "Human-readable agent name (e.g. 'Example Coder')."
},
"note": {
"type": "string",
"description": "Optional context: what the agent does, or why it's worth reviewing."
},
"email": {
"type": "string",
"description": "Optional contact email for the submitter — if you want a reply or a heads-up when the review publishes. Providing it always flags the submission to the Hlido team."
}
}
}
subscribe
subscribe
Preview — Wave 3 will add persistent webhook + RSS subscriptions. For now this returns the agent's current state plus advisory polling instructions (RSS at /changelog/feed.xml or polling /data/attestations/{slug}.json). Use this to register interest in being notified when a slug's verdict changes.
subscribe(slug: string, channel?: string)
// inputSchema
{
"type": "object",
"required": [
"slug"
],
"properties": {
"slug": {
"type": "string",
"description": "The Hlido slug to subscribe to (e.g. 'cursor', 'aider')"
},
"channel": {
"enum": [
"rss",
"json",
"webhook"
],
"type": "string",
"description": "Preferred notification channel. webhook is advisory only until Wave 3 ships."
}
}
}
trust_check
trust_check
The core Hlido trust query: is a specific AI agent trustworthy? Given one agent (by Hlido slug or product/homepage URL) it returns the independent Laddoo trust score (0-100), tier (VITAL/STEADY/FADING/FLATLINE), a one-line verdict, a claim-verification summary, and any known incidents. Call this FIRST — before delegating to, installing, or relying on another agent — to get a fast trust read. Returns no_review_found if the agent isn't in Hlido's corpus (then call request_quick_audit). For the full claim-by-claim evidence, follow up with get_scorecard.
trust_check(use_case?: string, agent_or_url: string)
// inputSchema
{
"type": "object",
"required": [
"agent_or_url"
],
"properties": {
"use_case": {
"type": "string",
"description": "Optional. The task you're considering this agent for (e.g. 'multi-file TypeScript refactor'); tailors the verdict to that use case when provided."
},
"agent_or_url": {
"type": "string",
"description": "The agent to check: either its Hlido slug (e.g. 'aider', 'cursor') or its product/homepage URL (e.g. 'https://cursor.com'). A URL is matched to the closest reviewed agent."
}
}
}
verify_claim
verify_claim
Fact-check one specific marketing or capability claim about an agent against Hlido's independent testing. Returns Hlido's verdict (PASS/FAIL/PARTIAL/UNKNOWN) with a quoted evidence snippet and its source surface — or an honest null when that exact claim wasn't tested (absence of evidence, not proof). Use this to validate a vendor's specific promise before you rely on it.
verify_claim(agent: string, claim: string)
// inputSchema
{
"type": "object",
"required": [
"agent",
"claim"
],
"properties": {
"agent": {
"type": "string",
"description": "The agent's Hlido slug or product URL (e.g. 'cursor')."
},
"claim": {
"type": "string",
"description": "The specific claim to verify, in plain language (e.g. 'works offline' or 'SOC 2 compliant')."
}
}
}
verify_transparency
verify_transparency
Check any AI agent's EU AI Act Article-50 transparency posture — including agents Hlido has NOT reviewed yet. Returns two clearly separated layers: (1) Hlido's independent register verdict when the agent is in our reviewed corpus, and (2) a live public-surface probe of the Article-50 signals (AI-interaction disclosure, machine-readable marking/provenance, deepfake/synthetic labelling, detection tool). Use before adopting or delegating to a tool ahead of the 2026-08-02 transparency obligations. The live probe is a first-pass surface read, NOT a compliance determination and NOT legal advice; an undetected signal means 'not discoverable on the public surface', not 'non-compliant'. Unreviewed agents are automatically queued for a full independent review.
verify_transparency(url: string)
// inputSchema
{
"type": "object",
"required": [
"url"
],
"properties": {
"url": {
"type": "string",
"description": "The agent's product or homepage URL (e.g. 'https://example.com'), or its Hlido slug."
}
}
}
Вопросы, новые серверы, обсуждение MCP
t.me/rusmcp · t.me/RusMcp_bot