Cookbook·News and media·indianexpress.com

Scrape section pages on indianexpress.com

The call that returns section pages from indianexpress.com today, with the parameters, patterns and failure modes read off the request log since 14 Aug 2026. Last tested 6 Sep 2026.

99.1% success · Aug 2026No browser neededStandard price · 1 credit

About indianexpress.com

indianexpress.com describes itself as “Get the latest news today, breaking news, and top news headlines updates from India and around the world. Stay updated on politics, business, sports, entertainment, and more with The Indian Express”.

Its home page declares 4 languages: en-in, en, en-us, en-ca.

Its sitemaps list 2,280 URLs across 19 files (news-sitemap, video-sitemap, category-sitemap, allsitemap, aboutsitemap1, aboutsitemap2). 418 of the 1,807 dated entries changed in the last 30 days, the newest on 14 Aug 2026.

Its robots.txt names 11 sitemaps and disallows 12 paths, among them /wp-admin/, /dummy/, /homepage/.

Measured on Crawlbase

Success rate over every request from every account in August 2026: 99.1%. Patterns, countries and failures come from the request log since 14 Aug 2026; 4.2% of successful calls used the JavaScript token and the median answer took 3.0s.

Among the 74 news and media sites in the Cookbook, indianexpress.com ranks 35 of 74 by success rate; the category median is 98.9%. Its median answer of 3.0s is slower than the category median of 2.3s. 58 of the 74 accept the plain token, this one among them.

The call

Article pages are the shape the accounts fetch most here, with the country and token that succeed most often. Replace the placeholder path with a real article URL from the site.

Run in Playground
curl "https://api.crawlbase.com/?token=YOUR_TOKEN&url=https%3A%2F%2Fwww.indianexpress.com%2Fsection%2Fexample-page"

Parameters that matter

ParameterSet it toWhy
countryleave unsetMost successful calls set no country (98.2%) and succeed at 99.3%.
javascriptleave unsetPlain-token calls succeed at 99.2% and make up 99.7% of the traffic. A browser costs more and, on this site, buys nothing.

URL patterns accounts fetch

PatternShare of successesJavaScript tokenCountryQuery parameters
/section/{slug}98.5%0.0%noneref

Fields you get

Structured data Crawlbase found on indianexpress.com, by page type. These are the schema.org types in the page, the fields your parser can read straight from the JSON-LD block:

  • Home: WebPage, ViewAction, NewsMediaOrganization, SiteNavigationElement, WebSite
  • Detail page: NewsMediaOrganization, NewsArticle, VideoObject, SiteNavigationElement, WebSite
  • Listing: WebPage, ViewAction, BreadcrumbList, ItemList, NewsMediaOrganization, SiteNavigationElement, WebSite
  • Other: WebPage, ViewAction, BreadcrumbList, ItemList, NewsMediaOrganization, SiteNavigationElement, WebSite

What breaks, and the fix

Of the failures since 14 Aug 2026, the site answered:

403: 76.0% of failures A 403 is the site refusing the exit. Set the country the successful calls use and retry once; a second 403 means the URL itself is protected.
0: 12.0% of failures A status of 0 is a request that never got an answer inside the timeout. Retry once with the JavaScript token if the pattern needs a browser.

Run it on a schedule

News moves by the hour. Poll the section or the sitemap for new links, then push the article URLs to the Crawler with a callback so the fetches are paced. The median answer on this site is 3.0s.

Questions

Do I need a browser to scrape indianexpress.com?

No. 99.2% of plain-token calls succeed and accounts send 99.7% of their traffic without a browser.

How much does a indianexpress.com page cost on Crawlbase?

1 credit without a browser. indianexpress.com is in the standard tier.

Which country should I set for indianexpress.com?

None is needed: most successful calls leave it unset.

Other pages on indianexpress.com

Pages indianexpress.com lists in its own sitemaps, fetched by Crawlbase on 6 Sep 2026 and verified in 2 consecutive runs. Each row is the cheapest call that returned the page with its content.

PageExampleTokenCountryMedian
Home
/
Structured data: WebPage, ViewAction, NewsMediaOrganization, SiteNavigationElement, WebSite
indianexpress.com/Plainnone5.6s
Detail page
/videos/{slug}/{slug}
Structured data: NewsMediaOrganization, NewsArticle, VideoObject, SiteNavigationElement, WebSite
indianexpress.com/videos/news-video/hemant-sorens-first-reaction-to-jharkhand-student-protests/Plainnone7.1s
Listing
/{slug}/{slug}/{slug}
Structured data: WebPage, ViewAction, BreadcrumbList, ItemList, NewsMediaOrganization, SiteNavigationElement, WebSite
nonePlainnone6.1s
Other
/{slug}/{slug}
Structured data: WebPage, ViewAction, BreadcrumbList, ItemList, NewsMediaOrganization, SiteNavigationElement, WebSite
nonePlainnone5.3s
Site names identify the pages a recipe fetches. Crawlbase is not affiliated with any site listed. Use recipes within the acceptable use policy.Measured numbers from August 2026, refreshed monthly