Scrape articles on yahoo.com
The call that returns articles from yahoo.com today, with the parameters, patterns and failure modes read off the request log since 14 Aug 2026. Last tested 6 Sep 2026.
About yahoo.com
yahoo.com describes itself as “Latest news coverage, email, free stock quotes, live scores and video are just the beginning. Discover more every day at Yahoo!”.
Its sitemaps list 109,110 URLs across 19 files (index).
Its robots.txt names 3 sitemaps and disallows 12 paths, among them /p/, /r/, /bin/.
Measured on Crawlbase
Success rate over every request from every account in August 2026: 99.6%. Patterns, countries and failures come from the request log since 14 Aug 2026; 1.5% of successful calls used the JavaScript token and the median answer took 1.0s.
Among the 74 news and media sites in the Cookbook, yahoo.com ranks 26 of 74 by success rate; the category median is 98.9%. Its median answer of 1.0s is faster than the category median of 2.3s. 58 of the 74 accept the plain token, this one among them.
The call
Article pages are the shape the accounts fetch most here, with the country and token that succeed most often. Replace the placeholder path with a real article URL from the site.
curl "https://api.crawlbase.com/?token=YOUR_TOKEN&url=https%3A%2F%2Fwww.yahoo.com%2Fexample-section%2Fexample-page%2Fexample-page"from crawlbase import CrawlingAPI api = CrawlingAPI({'token': 'YOUR_TOKEN'}) r = api.get('https://www.yahoo.com/example-section/example-page/example-page') html = r['body']
const { CrawlingAPI } = require('crawlbase'); const api = new CrawlingAPI({ token: 'YOUR_TOKEN' }); const r = await api.get('https://www.yahoo.com/example-section/example-page/example-page');
Parameters that matter
| Parameter | Set it to | Why |
|---|---|---|
country | leave unset | Most successful calls set no country (98.6%) and succeed at 99.8%. |
javascript | leave unset | Plain-token calls succeed at 99.8% and make up 98.2% of the traffic. A browser costs more and, on this site, buys nothing. |
URL patterns accounts fetch
| Pattern | Share of successes | JavaScript token | Country | Query parameters |
|---|---|---|---|---|
/{slug}/{slug}/{slug} | 33.0% | 0.0% | none | period1, period2, interval |
/quote/{slug} | 19.5% | 0.0% | none | none |
Fields you get
Structured data Crawlbase found on yahoo.com, by page type. These are the schema.org types in the page, the fields your parser can read straight from the JSON-LD block:
- Home:
WebSite, NewsMediaOrganization
What breaks, and the fix
Of the failures since 14 Aug 2026, the site answered:
Run it on a schedule
News moves by the hour. Poll the section or the sitemap for new links, then push the article URLs to the Crawler with a callback so the fetches are paced. The median answer on this site is 1.0s.
Questions
Do I need a browser to scrape yahoo.com?
No. 99.8% of plain-token calls succeed and accounts send 98.2% of their traffic without a browser.
How much does a yahoo.com page cost on Crawlbase?
1 credit without a browser. yahoo.com is in the standard tier.
Which country should I set for yahoo.com?
None is needed: most successful calls leave it unset.
Other pages on yahoo.com
Pages yahoo.com lists in its own sitemaps, fetched by Crawlbase on 6 Sep 2026 and verified in 2 consecutive runs. Each row is the cheapest call that returned the page with its content.
| Page | Example | Token | Country | Median |
|---|---|---|---|---|
Home/Structured data: WebSite, NewsMediaOrganization | yahoo.com/ | Plain | none | 2.3s |
Article/news/{slug}/{slug} | yahoo.com/news/weather/austria/tyrol/silz-551226 | JavaScript | none | 6.4s |