Scrape justice.gov
The call that returns the homepage from justice.gov today, with the parameters, patterns and failure modes read off the request log since 14 Aug 2026. Last tested 6 Sep 2026.
About justice.gov
justice.gov describes itself as “Official website of the U.S. Department of Justice (DOJ). DOJ’s mission is to enforce the law and defend the interests of the United States according to the law; to ensure public safety against threats foreign and domestic; to provide federal leadership in preventing and controlling crime; to seek…”.
Its sitemaps list 38,000 URLs across 19 files. 158 of the 38,000 dated entries changed in the last 30 days, the newest on 2 Sep 2026.
Its robots.txt names 1 sitemap and disallows 12 paths, among them /core/, /profiles/, /admin/. It asks crawlers to wait 10s between requests.
Measured on Crawlbase
Success rate over every request from every account in August 2026: 97.2%. Patterns, countries and failures come from the request log since 14 Aug 2026; 57.6% of successful calls used the JavaScript token and the median answer took 9.8s.
Among the 11 public data sites in the Cookbook, justice.gov ranks 6 of 11 by success rate; the category median is 97.2%. Its median answer of 9.8s is slower than the category median of 2.3s. 5 of the 11 accept the plain token; this one does on some pages only.
The call
The most common shape of URL the accounts fetch, with the country and token that succeed most often. Replace the placeholder path with a real one from the site.
curl "https://api.crawlbase.com/?token=YOUR_TOKEN&url=https%3A%2F%2Fwww.justice.gov%2F"from crawlbase import CrawlingAPI api = CrawlingAPI({'token': 'YOUR_TOKEN'}) r = api.get('https://www.justice.gov/') html = r['body']
const { CrawlingAPI } = require('crawlbase'); const api = new CrawlingAPI({ token: 'YOUR_TOKEN' }); const r = await api.get('https://www.justice.gov/');
Parameters that matter
| Parameter | Set it to | Why |
|---|---|---|
country | leave unset | Most successful calls set no country (100%) and succeed at 97.5%. |
javascript | leave unset | Plain-token calls succeed at 98.2%, JavaScript-token calls at 97.0%. Both work; start without the browser and switch per URL pattern where the table below shows it. |
Fields you get
Fields depend on the page type. Product pages usually carry an ld+json block with name, price, currency and availability, and article pages a headline and body; check the first response you get.
What breaks, and the fix
Of the failures since 14 Aug 2026, the site answered:
Run it on a schedule
The median answer on this site is 9.8s. Push a URL list to the Crawler with a callback for anything larger than a few thousand pages a day; it paces the calls and retries what the site refuses.
Questions
Do I need a browser to scrape justice.gov?
It depends on the page. Plain-token calls succeed at 98.2%, JavaScript-token calls at 97.0%; the URL pattern table shows which paths accounts fetch with a browser.
How much does a justice.gov page cost on Crawlbase?
1 credit without a browser. justice.gov is in the standard tier.
Which country should I set for justice.gov?
None is needed: most successful calls leave it unset.
Other pages on justice.gov
Pages justice.gov lists in its own sitemaps, fetched by Crawlbase on 6 Sep 2026 and verified in 2 consecutive runs. Each row is the cheapest call that returned the page with its content.
| Page | Example | Token | Country | Median |
|---|---|---|---|---|
Home/ | justice.gov/ | Plain | none | 1.8s |
Article/{slug}/{slug}/blog | none | JavaScript | none | 18.7s |
Other/{slug}/{slug}/{slug} | none | Plain | none | 6.7s |