crawlfox (Elixir)
Official Elixir SDK for the CrawlFox API: scrape, search (Google, Bing, or DuckDuckGo), batch scrape, map site links, crawl multi-page jobs, and logs.
Get a key from the dashboard. Set CRAWLFOX_API_KEY.
Hex once published:
{:crawlfox, "~> 0.2.0"}
Until Hex lists it:
{:crawlfox,
git: "https://github.com/Automote-LLC/crawlfox-integrations.git",
sparse: "packages/elixir"}
client = Crawlfox.new()
page = Crawlfox.scrape(client, "https://example.com", %{
"formats" => ["markdown", "links"]
})
hits = Crawlfox.search(client, "rust async tutorial", %{
"engine" => "google",
"num" => 10
})
batch = Crawlfox.batch(client, ["https://example.com/", "https://example.org/"], %{
"formats" => ["markdown"]
})
Formats: markdown, html, rawHtml, json, links, images, emails. CSS selectors live under jsonOptions when you request json.
Map
POST /v1/map. One credit per call.
mapped = Crawlfox.map(client, "https://example.com/", %{limit: 10})
Crawl
job = Crawlfox.crawl(client, "https://example.com/", %{
limit: 3,
scrapeOptions: %{formats: ["markdown"]}
})
status = Crawlfox.get_crawl(client, job["id"])
done = Crawlfox.wait_for_crawl(client, job["id"])
Crawlfox.cancel_crawl(client, job["id"])
errors = Crawlfox.get_crawl_errors(client, job["id"])
Credits: scrape 1 per page, search 1 per 10 requested results, map 1 per call, crawl 1 per page per format.
mix deps.get && mix test
Docs: https://docs.crawlfox.io