Перейти к содержанию

Firecrawl

Поиск в интернете, извлечение содержимого страниц и разбор документов

Поиск и данные из интернетаавтор: Firecrawl · добавлен каталогом✓ Проверен модератором
3 инструмента
подключить к ассистенту

Добавьте сервер в «Мои MCP» — получите адрес подключения для Claude и ChatGPT.

Добавить в «Мои MCP»

Нужно войти или зарегистрироваться — вернём на эту страницу.

/about

Описание

Firecrawl — сервис для получения данных из интернета в удобном для ИИ виде. Подойдёт тем, кому ассистент должен читать сайты, искать свежую информацию и разбирать документы. Три инструмента: загрузка одной страницы с выдачей в Markdown, HTML, списком ссылок, скриншотом, данными о брендинге, точечным ответом или JSON по заданной схеме; поиск по вебу, новостям и изображениям с релевантными выдержками; разбор документа — HTML, PDF, Word, RTF, OpenDocument, таблиц — в Markdown, краткое содержание или структурированные данные. Для работы нужен аккаунт Firecrawl.

/tools

Инструменты · 3

из ответа tools/list

Firecrawl file parsing

firecrawl_parse

Parse one supported document into markdown, HTML, links, summary, targeted answers, or JSON matching a schema. Supported inputs include common HTML, PDF, Word, RTF, OpenDocument, and spreadsheet files; PDF parsing can be bounded with `pdfOptions.maxPages`. Local MCP reads `filePath` from the server filesystem. Hosted MCP uses two calls: first provide `filePath` to receive upload instructions, upload locally, then call again with the returned `uploadRef`; do not send both fields together. Remote web URLs belong in `firecrawl_scrape`. Set `redactPII` to request redaction of personally identifiable information in the returned content. `zeroDataRetention` requires an eligible authenticated account; omit it for anonymous keyless use. Returns upload instructions for hosted phase one or parsed document content for the final call. Authenticated final responses can include a `data.metadata.scrapeId` for optional parse feedback.

firecrawl_parse(proxy?: string, maxAge?: number, formats?: array, parsers?: array, filePath?: string, redactPII?: boolean, uploadRef?: string, pdfOptions?: object, contentType?: string, excludeTags?: array, includeTags?: array, jsonOptions?: object, queryOptions?: object, storeInCache?: boolean, onlyMainContent?: boolean, declaredSizeBytes?: integer, zeroDataRetention?: boolean, removeBase64Images?: boolean, skipTlsVerification?: boolean)

Firecrawl scrape

firecrawl_scrape

Scrape one URL and return its content: markdown by default, or HTML, links, screenshots, branding data, a targeted answer, or JSON matching a supplied schema. Use it when the request identifies a page and needs its content or defined fields. Use `firecrawl_search` when additional web sources are needed; on an authenticated session, `firecrawl_map` lists a site's URLs and `firecrawl_crawl` collects a set of pages. Firecrawl may serve recently indexed content; set `maxAge: 0` for a live fetch or a smaller `maxAge` to bound staleness. A successful response does not by itself confirm the page is still current. A named browser profile loads saved session data without saving changes to it. Authenticated responses can include a `metadata.scrapeId` for optional scrape feedback. On an authenticated session with Alexandria access, `firecrawl_search` with `sources` unset and `firecrawl_find_tools` can discover providers for the same fields across several pages; a matching provider returns typed records in one call. Keyless sessions have no provider matches. Alexandria mode, on an authenticated session with Alexandria access: `alexandria` selects catalogued capability execution and is mutually exclusive with `url`.

firecrawl_scrape(url?: string, proxy?: string, maxAge?: number, mobile?: boolean, formats?: array, parsers?: array, profile?: object, timeout?: integer, waitFor?: number, location?: object, lockdown?: boolean, redactPII?: boolean, requestId?: string, alexandria?: any, pdfOptions?: object, toolDetail?: string, domainTools?: boolean, excludeTags?: array, includeTags?: array, jsonOptions?: object, queryOptions?: object, storeInCache?: boolean, onlyMainContent?: boolean, screenshotOptions?: object, zeroDataRetention?: boolean, removeBase64Images?: boolean, skipTlsVerification?: boolean)

Firecrawl web search

firecrawl_search

Search web, news, or image sources and return ranked results with query-relevant highlights. Each web result is a title, URL, and description; use `firecrawl_scrape` on a result URL when the excerpt is not enough. Authenticated search also returns matching Alexandria data providers in data.tools (companies, people, jobs, finance and filings, public records and government spending, real estate, places and restaurants, retail and prices, package registries and developer data, news, research, and more). Prefer a provider over scraping pages when the task needs the same fields across several entities, exact figures or timestamps, provenance, or many records; use web results when they already answer the question. A search with sources: ["web"] omits semantic provider discovery; domainTools: true can still return website-matched tools. Web-only results use domainTools: false. On an authenticated session, tool matches describe available capabilities; `firecrawl_find_tools` returns their contracts and `firecrawl_scrape` with an `alexandria` body executes a selected capability. Keyless sessions get no Alexandria matches in data.tools. For a programming question, add `categories: ["developer"]`; its hits return in `data.web` with `category: "developer"`. For a legal or regulatory question, `categories: ["gov"]` returns hits in `data.web` with `category: "gov"` and cannot be combined with other categories; authenticated sessions also have `firecrawl_gov_search` as the dedicated tool. `categories: ["research"]` restricts web results to research-affiliated websites; the `firecrawl_research_*` tools are a separate surface over paper abstracts and full text (PubMed, bioRxiv, medRxiv, arXiv). Query operators, domain filters, `categories`, `toolDetail` and `scrapeOptions` are described on their parameters. Returns source-type result groups and usage metadata. Authenticated responses can include an `id` for optional search feedback.

firecrawl_search(tbs?: string, limit?: integer, query: string, filter?: string, sources?: array, location?: string, objective?: string, categories?: array, enterprise?: array, highlights?: boolean, toolDetail?: string, clientModel?: string, domainTools?: boolean, scrapeOptions?: object, excludeDomains?: array, includeDomains?: array)

Вопросы, новые серверы, обсуждение MCP

t.me/rusmcp · t.me/RusMcp_bot