crawlfox (Elixir)

Official Elixir SDK for the CrawlFox API: scrape, search (Google, Bing, or DuckDuckGo), batch scrape, map site links, crawl multi-page jobs, and logs.

Get a key from the dashboard. Set CRAWLFOX_API_KEY.

Hex once published:

{:crawlfox, "~> 0.2.0"}

Until Hex lists it:

{:crawlfox,
git: "https://github.com/Automote-LLC/crawlfox-integrations.git",
sparse: "packages/elixir"}
client = Crawlfox.new()
page = Crawlfox.scrape(client, "https://example.com", %{
"formats" => ["markdown", "links"]
})
hits = Crawlfox.search(client, "rust async tutorial", %{
"engine" => "google",
"num" => 10
})
batch = Crawlfox.batch(client, ["https://example.com/", "https://example.org/"], %{
"formats" => ["markdown"]
})

Formats: markdown, html, rawHtml, json, links, images, emails. CSS selectors live under jsonOptions when you request json.

Map

POST /v1/map. One credit per call.

mapped = Crawlfox.map(client, "https://example.com/", %{limit: 10})

Crawl

job = Crawlfox.crawl(client, "https://example.com/", %{
limit: 3,
scrapeOptions: %{formats: ["markdown"]}
})
status = Crawlfox.get_crawl(client, job["id"])
done = Crawlfox.wait_for_crawl(client, job["id"])
Crawlfox.cancel_crawl(client, job["id"])
errors = Crawlfox.get_crawl_errors(client, job["id"])

Credits: scrape 1 per page, search 1 per 10 requested results, map 1 per call, crawl 1 per page per format.

mix deps.get && mix test

Docs: https://docs.crawlfox.io