Firecrawl
Поиск в интернете, извлечение содержимого страниц и разбор документов
Добавьте сервер в «Мои MCP» — получите адрес подключения для Claude и ChatGPT.
Добавить в «Мои MCP»Нужно войти или зарегистрироваться — вернём на эту страницу.
Описание
Firecrawl — сервис для получения данных из интернета в удобном для ИИ виде. Подойдёт тем, кому ассистент должен читать сайты, искать свежую информацию и разбирать документы. Три инструмента: загрузка одной страницы с выдачей в Markdown, HTML, списком ссылок, скриншотом, данными о брендинге, точечным ответом или JSON по заданной схеме; поиск по вебу, новостям и изображениям с релевантными выдержками; разбор документа — HTML, PDF, Word, RTF, OpenDocument, таблиц — в Markdown, краткое содержание или структурированные данные. Для работы нужен аккаунт Firecrawl.
Инструменты · 3
из ответа tools/list
Firecrawl file parsing
firecrawl_parse
Firecrawl file parsing
firecrawl_parse Parse one supported document into markdown, HTML, links, summary, targeted answers, or JSON matching a schema. Supported inputs include common HTML, PDF, Word, RTF, OpenDocument, and spreadsheet files; PDF parsing can be bounded with `pdfOptions.maxPages`. Local MCP reads `filePath` from the server filesystem. Hosted MCP uses two calls: first provide `filePath` to receive upload instructions, upload locally, then call again with the returned `uploadRef`; do not send both fields together. Remote web URLs belong in `firecrawl_scrape`. Set `redactPII` to request redaction of personally identifiable information in the returned content. `zeroDataRetention` requires an eligible authenticated account; omit it for anonymous keyless use. Returns upload instructions for hosted phase one or parsed document content for the final call. Authenticated final responses can include a `data.metadata.scrapeId` for optional parse feedback.
firecrawl_parse(proxy?: string, maxAge?: number, formats?: array, parsers?: array, filePath?: string, redactPII?: boolean, uploadRef?: string, pdfOptions?: object, contentType?: string, excludeTags?: array, includeTags?: array, jsonOptions?: object, queryOptions?: object, storeInCache?: boolean, onlyMainContent?: boolean, declaredSizeBytes?: integer, zeroDataRetention?: boolean, removeBase64Images?: boolean, skipTlsVerification?: boolean)
// inputSchema
{
"type": "object",
"$schema": "http://json-schema.org/draft-07/schema#",
"properties": {
"proxy": {
"enum": [
"basic",
"auto"
],
"type": "string"
},
"maxAge": {
"type": "number",
"description": "Ignored: parse never reuses or stores indexed content."
},
"formats": {
"type": "array",
"items": {
"enum": [
"markdown",
"html",
"rawHtml",
"links",
"summary",
"json",
"query"
],
"type": "string"
}
},
"parsers": {
"type": "array",
"items": {
"enum": [
"pdf"
],
"type": "string"
}
},
"filePath": {
"type": "string",
"minLength": 1,
"description": "Phase 1 only: path to the local file on the caller/harness machine. Hosted MCP will not read or stat this path; it is used only to produce upload instructions."
},
"redactPII": {
"type": "boolean"
},
"uploadRef": {
"type": "string",
"minLength": 1,
"description": "Phase 2 only: short-lived upload reference returned by phase 1 after the local PUT upload completes."
},
"pdfOptions": {
"type": "object",
"properties": {
"maxPages": {
"type": "integer",
"maximum": 10000,
"minimum": 1
}
},
"additionalProperties": false
},
"contentType": {
"type": "string",
"description": "Phase 1 MIME type override. If omitted, the server infers it from the file extension without reading the file."
},
"excludeTags": {
"type": "array",
"items": {
"type": "string"
}
},
"includeTags": {
"type": "array",
"items": {
"type": "string"
}
},
"jsonOptions": {
"type": "object",
"properties": {
"prompt": {
"type": "string"
},
"schema": {
"type": "object",
"propertyNames": {
"type": "string"
},
"additionalProperties": false
}
},
"additionalProperties": false
},
"queryOptions": {
"type": "object",
"required": [
"prompt"
],
"properties": {
"mode": {
"enum": [
"directQuote",
"freeform"
],
"type": "string",
"default": "freeform"
},
"prompt": {
"type": "string",
"maxLength": 10000
}
},
"additionalProperties": false
},
"storeInCache": {
"type": "boolean"
},
"onlyMainContent": {
"type": "boolean"
},
"declaredSizeBytes": {
"type": "integer",
"maximum": 9007199254740991,
"description": "Optional phase 1 size declaration. Hosted MCP does not stat the file; provide this only if the caller already knows it.",
"exclusiveMinimum": 0
},
"zeroDataRetention": {
"type": "boolean"
},
"removeBase64Images": {
"type": "boolean"
},
"skipTlsVerification": {
"type": "boolean"
}
},
"additionalProperties": false
}
Firecrawl scrape
firecrawl_scrape
Firecrawl scrape
firecrawl_scrape Scrape one URL and return its content: markdown by default, or HTML, links, screenshots, branding data, a targeted answer, or JSON matching a supplied schema. Use it when the request identifies a page and needs its content or defined fields. Use `firecrawl_search` when additional web sources are needed; on an authenticated session, `firecrawl_map` lists a site's URLs and `firecrawl_crawl` collects a set of pages. Firecrawl may serve recently indexed content; set `maxAge: 0` for a live fetch or a smaller `maxAge` to bound staleness. A successful response does not by itself confirm the page is still current. A named browser profile loads saved session data without saving changes to it. Authenticated responses can include a `metadata.scrapeId` for optional scrape feedback. On an authenticated session with Alexandria access, `firecrawl_search` with `sources` unset and `firecrawl_find_tools` can discover providers for the same fields across several pages; a matching provider returns typed records in one call. Keyless sessions have no provider matches. Alexandria mode, on an authenticated session with Alexandria access: `alexandria` selects catalogued capability execution and is mutually exclusive with `url`.
firecrawl_scrape(url?: string, proxy?: string, maxAge?: number, mobile?: boolean, formats?: array, parsers?: array, profile?: object, timeout?: integer, waitFor?: number, location?: object, lockdown?: boolean, redactPII?: boolean, requestId?: string, alexandria?: any, pdfOptions?: object, toolDetail?: string, domainTools?: boolean, excludeTags?: array, includeTags?: array, jsonOptions?: object, queryOptions?: object, storeInCache?: boolean, onlyMainContent?: boolean, screenshotOptions?: object, zeroDataRetention?: boolean, removeBase64Images?: boolean, skipTlsVerification?: boolean)
// inputSchema
{
"type": "object",
"$schema": "http://json-schema.org/draft-07/schema#",
"properties": {
"url": {
"type": "string",
"format": "uri"
},
"proxy": {
"enum": [
"basic",
"stealth",
"enhanced",
"auto"
],
"type": "string"
},
"maxAge": {
"type": "number"
},
"mobile": {
"type": "boolean"
},
"formats": {
"type": "array",
"items": {
"enum": [
"markdown",
"html",
"rawHtml",
"screenshot",
"links",
"summary",
"changeTracking",
"branding",
"json",
"query",
"audio"
],
"type": "string"
}
},
"parsers": {
"type": "array",
"items": {
"enum": [
"pdf"
],
"type": "string"
}
},
"profile": {
"type": "object",
"required": [
"name"
],
"properties": {
"name": {
"type": "string"
}
},
"description": "Loads a saved browser profile without saving changes to it.",
"additionalProperties": false
},
"timeout": {
"type": "integer",
"maximum": 9007199254740991,
"description": "Execution timeout in milliseconds.",
"exclusiveMinimum": 0
},
"waitFor": {
"type": "number"
},
"location": {
"type": "object",
"properties": {
"country": {
"type": "string"
},
"languages": {
"type": "array",
"items": {
"type": "string"
}
}
},
"additionalProperties": false
},
"lockdown": {
"type": "boolean"
},
"redactPII": {
"type": "boolean"
},
"requestId": {
"type": "string",
"pattern": "^[A-Za-z0-9._:-]{1,128}$",
"description": "Idempotency key bound to one Alexandria execution payload. Generated when omitted and returned with the result."
},
"alexandria": {
"anyOf": [
{
"type": "object",
"required": [
"provider",
"capability"
],
"properties": {
"options": {
"type": "object",
"description": "Capability options as declared by its contract.",
"propertyNames": {
"type": "string"
},
"additionalProperties": {}
},
"version": {
"type": "string",
"maxLength": 128,
"minLength": 1,
"description": "Optional published workflow version. Omit to use the latest version."
},
"provider": {
"type": "string",
"minLength": 1,
"description": "Provider slug, e.g. \"fred\"."
},
"capability": {
"type": "string",
"minLength": 1,
"description": "Capability address as returned by search or discover, e.g. \"series/observations\"."
}
}
},
{
"type": "array",
"items": {
"type": "object",
"required": [
"provider",
"capability"
],
"properties": {
"options": {
"type": "object",
"description": "Capability options as declared by its contract.",
"propertyNames": {
"type": "string"
},
"additionalProperties": {}
},
"version": {
"type": "string",
"maxLength": 128,
"minLength": 1,
"description": "Optional published workflow version. Omit to use the latest version."
},
"provider": {
"type": "string",
"minLength": 1,
"description": "Provider slug, e.g. \"fred\"."
},
"capability": {
"type": "string",
"minLength": 1,
"description": "Capability address as returned by search or discover, e.g. \"series/observations\"."
}
}
},
"maxItems": 10,
"minItems": 1
}
],
"description": "Catalogued Alexandria capability invocation, mutually exclusive with url. One {provider, capability, options} object or an array of 1-10, with contracts available through firecrawl_search or firecrawl_find_tools. Each call may include version to pin a published workflow; omitting it uses latest. Only requestId and timeout are supported alongside alexandria. The selected contract marks required inputs and any requiresOneOf groups (at least one member per group); it may include example.request/example.response and response.key (which may differ from records). Where pagination is declared, its fields govern paging with the same filters; catalogue next is separate from provider pagination. Returns per-capability results in data.alexandria with data, records, or an error with a code; individual capabilities can fail even when the outer response succeeds. Requires an API key on a team with Alexandria enabled. Some providers require accepted terms; blocked requests return the applicable requirements."
},
"pdfOptions": {
"type": "object",
"properties": {
"maxPages": {
"type": "integer",
"maximum": 10000,
"minimum": 1
}
},
"additionalProperties": false
},
"toolDetail": {
"enum": [
"compact",
"summary",
"full"
],
"type": "string",
"description": "URL mode only: domain discovery detail, summary by default; compact returns provider/capability/description, full includes contracts. Ignored with alexandria."
},
"domainTools": {
"type": "boolean",
"description": "URL mode only: include domain-matched Alexandria tools for the page in tools on the returned document. Ignored with alexandria."
},
"excludeTags": {
"type": "array",
"items": {
"type": "string"
}
},
"includeTags": {
"type": "array",
"items": {
"type": "string"
}
},
"jsonOptions": {
"type": "object",
"properties": {
"prompt": {
"type": "string"
},
"schema": {
"type": "object",
"propertyNames": {
"type": "string"
},
"additionalProperties": false
}
},
"additionalProperties": false
},
"queryOptions": {
"type": "object",
"required": [
"prompt"
],
"properties": {
"mode": {
"enum": [
"directQuote",
"freeform"
],
"type": "string",
"default": "freeform"
},
"prompt": {
"type": "string",
"maxLength": 10000
}
},
"additionalProperties": false
},
"storeInCache": {
"type": "boolean"
},
"onlyMainContent": {
"type": "boolean"
},
"screenshotOptions": {
"type": "object",
"properties": {
"quality": {
"type": "number"
},
"fullPage": {
"type": "boolean"
},
"viewport": {
"type": "object",
"required": [
"width",
"height"
],
"properties": {
"width": {
"type": "number"
},
"height": {
"type": "number"
}
},
"additionalProperties": false
}
},
"additionalProperties": false
},
"zeroDataRetention": {
"type": "boolean"
},
"removeBase64Images": {
"type": "boolean"
},
"skipTlsVerification": {
"type": "boolean"
}
},
"additionalProperties": false
}
Firecrawl web search
firecrawl_search
Firecrawl web search
firecrawl_search Search web, news, or image sources and return ranked results with query-relevant highlights. Each web result is a title, URL, and description; use `firecrawl_scrape` on a result URL when the excerpt is not enough. Authenticated search also returns matching Alexandria data providers in data.tools (companies, people, jobs, finance and filings, public records and government spending, real estate, places and restaurants, retail and prices, package registries and developer data, news, research, and more). Prefer a provider over scraping pages when the task needs the same fields across several entities, exact figures or timestamps, provenance, or many records; use web results when they already answer the question. A search with sources: ["web"] omits semantic provider discovery; domainTools: true can still return website-matched tools. Web-only results use domainTools: false. On an authenticated session, tool matches describe available capabilities; `firecrawl_find_tools` returns their contracts and `firecrawl_scrape` with an `alexandria` body executes a selected capability. Keyless sessions get no Alexandria matches in data.tools. For a programming question, add `categories: ["developer"]`; its hits return in `data.web` with `category: "developer"`. For a legal or regulatory question, `categories: ["gov"]` returns hits in `data.web` with `category: "gov"` and cannot be combined with other categories; authenticated sessions also have `firecrawl_gov_search` as the dedicated tool. `categories: ["research"]` restricts web results to research-affiliated websites; the `firecrawl_research_*` tools are a separate surface over paper abstracts and full text (PubMed, bioRxiv, medRxiv, arXiv). Query operators, domain filters, `categories`, `toolDetail` and `scrapeOptions` are described on their parameters. Returns source-type result groups and usage metadata. Authenticated responses can include an `id` for optional search feedback.
firecrawl_search(tbs?: string, limit?: integer, query: string, filter?: string, sources?: array, location?: string, objective?: string, categories?: array, enterprise?: array, highlights?: boolean, toolDetail?: string, clientModel?: string, domainTools?: boolean, scrapeOptions?: object, excludeDomains?: array, includeDomains?: array)
// inputSchema
{
"type": "object",
"$schema": "http://json-schema.org/draft-07/schema#",
"required": [
"query"
],
"properties": {
"tbs": {
"type": "string"
},
"limit": {
"type": "integer",
"maximum": 100,
"minimum": 1
},
"query": {
"type": "string",
"minLength": 1,
"description": "Query for web and semantic tool discovery. Operators include quoted phrases, `-term`, `site:host`, `inurl:term`, `intitle:term`, and `related:host`; the set is non-exhaustive. Catalogue browsing is available through firecrawl_find_tools."
},
"filter": {
"type": "string"
},
"sources": {
"type": "array",
"items": {
"anyOf": [
{
"enum": [
"web",
"images",
"news",
"alexandria",
"exchange"
],
"type": "string"
},
{
"type": "object",
"required": [
"type"
],
"properties": {
"type": {
"enum": [
"web",
"images",
"news",
"alexandria",
"exchange"
],
"type": "string"
}
},
"additionalProperties": false
}
]
},
"description": "Search sources; authenticated sessions default to web + alexandria, keyless sessions to web only. A search with sources: [\"web\"] omits semantic provider discovery; domainTools: true can still return website-matched tools. Web-only results use domainTools: false. Use [\"alexandria\"] alone for provider discovery without web results."
},
"location": {
"type": "string"
},
"objective": {
"type": "string",
"maxLength": 5000,
"minLength": 1,
"description": "Optional broader goal for this search, if known. Avoid sensitive information."
},
"categories": {
"type": "array",
"items": {
"enum": [
"research",
"pdf",
"developer",
"gov"
],
"type": "string"
},
"description": "Limit results to specific source types. `research` restricts ordinary web results to research-affiliated websites and returns page snippets, which is separate from the `firecrawl_research_*` tools that search paper abstracts and full text across biomedical (PubMed, bioRxiv, medRxiv) and arXiv literature; `pdf` searches PDF results; `developer` searches an index built for coding agents over public repositories, GitHub issues, merged pull requests, repository READMEs, and code documentation; `gov` searches the Government Index of US federal, state, and local legal and regulatory sources, the index behind firecrawl_gov_search, and cannot be combined with other categories. `developer` and `gov` return hits in `data.web` with `category` set to their name; the other categories also filter `data.web`."
},
"enterprise": {
"type": "array",
"items": {
"enum": [
"default",
"anon",
"zdr"
],
"type": "string"
}
},
"highlights": {
"type": "boolean",
"description": "Return query-relevant page excerpts for web and news results when available (default). Highlights appear in web `description` and news `snippet`; otherwise, original snippets are returned. Set to false to keep the original search snippets."
},
"toolDetail": {
"enum": [
"compact",
"summary",
"full"
],
"type": "string",
"description": "Compact by default. Compact returns only provider, capability and description; full includes contracts. Inspect selected compact tools with firecrawl_find_tools providers and capabilities."
},
"clientModel": {
"type": "string",
"maxLength": 128,
"minLength": 1,
"description": "Optional model identifier, if known."
},
"domainTools": {
"type": "boolean",
"description": "Include domain-matched tools for result URLs. Defaults to true when Alexandria is combined with web, news or images; semantic-only search leaves domain matching off."
},
"scrapeOptions": {
"type": "object",
"properties": {
"proxy": {
"enum": [
"basic",
"stealth",
"enhanced",
"auto"
],
"type": "string"
},
"maxAge": {
"type": "number"
},
"mobile": {
"type": "boolean"
},
"formats": {
"type": "array",
"items": {
"enum": [
"markdown",
"html",
"rawHtml",
"screenshot",
"links",
"summary",
"changeTracking",
"branding",
"json",
"query",
"audio"
],
"type": "string"
}
},
"parsers": {
"type": "array",
"items": {
"enum": [
"pdf"
],
"type": "string"
}
},
"profile": {
"type": "object",
"required": [
"name"
],
"properties": {
"name": {
"type": "string"
}
},
"description": "Loads a saved browser profile without saving changes to it.",
"additionalProperties": false
},
"waitFor": {
"type": "number"
},
"location": {
"type": "object",
"properties": {
"country": {
"type": "string"
},
"languages": {
"type": "array",
"items": {
"type": "string"
}
}
},
"additionalProperties": false
},
"lockdown": {
"type": "boolean"
},
"redactPII": {
"type": "boolean"
},
"pdfOptions": {
"type": "object",
"properties": {
"maxPages": {
"type": "integer",
"maximum": 10000,
"minimum": 1
}
},
"additionalProperties": false
},
"excludeTags": {
"type": "array",
"items": {
"type": "string"
}
},
"includeTags": {
"type": "array",
"items": {
"type": "string"
}
},
"jsonOptions": {
"type": "object",
"properties": {
"prompt": {
"type": "string"
},
"schema": {
"type": "object",
"propertyNames": {
"type": "string"
},
"additionalProperties": false
}
},
"additionalProperties": false
},
"queryOptions": {
"type": "object",
"required": [
"prompt"
],
"properties": {
"mode": {
"enum": [
"directQuote",
"freeform"
],
"type": "string",
"default": "freeform"
},
"prompt": {
"type": "string",
"maxLength": 10000
}
},
"additionalProperties": false
},
"storeInCache": {
"type": "boolean"
},
"onlyMainContent": {
"type": "boolean"
},
"screenshotOptions": {
"type": "object",
"properties": {
"quality": {
"type": "number"
},
"fullPage": {
"type": "boolean"
},
"viewport": {
"type": "object",
"required": [
"width",
"height"
],
"properties": {
"width": {
"type": "number"
},
"height": {
"type": "number"
}
},
"additionalProperties": false
}
},
"additionalProperties": false
},
"zeroDataRetention": {
"type": "boolean"
},
"removeBase64Images": {
"type": "boolean"
},
"skipTlsVerification": {
"type": "boolean"
}
},
"description": "Attach page content for web results in the same call. These fetches ignore maxAge, so use firecrawl_scrape when you need a live fetch. scrapeOptions fetches web pages, never Alexandria provider tools.",
"additionalProperties": false
},
"excludeDomains": {
"type": "array",
"items": {
"type": "string",
"pattern": "^(?:[a-z0-9](?:[a-z0-9-]{0,61}[a-z0-9])?\\.)+[a-z0-9][a-z0-9-]{0,61}[a-z0-9]$",
"maxLength": 253,
"minLength": 1
},
"description": "Hostnames to leave out of results. Mutually exclusive with includeDomains."
},
"includeDomains": {
"type": "array",
"items": {
"type": "string",
"pattern": "^(?:[a-z0-9](?:[a-z0-9-]{0,61}[a-z0-9])?\\.)+[a-z0-9][a-z0-9-]{0,61}[a-z0-9]$",
"maxLength": 253,
"minLength": 1
},
"description": "Hostnames to restrict results to. Mutually exclusive with excludeDomains."
}
},
"additionalProperties": false
}
Вопросы, новые серверы, обсуждение MCP
t.me/rusmcp · t.me/RusMcp_bot