News-articles datasets are ready-to-license structured news data — clean article text, headlines, publication dates, sources, extracted entities and sentiment scores — from news outlets. iWeb Data Scraping delivers them de-duplicated and enriched, so NLP teams, researchers and investors can work with news data without building collection, cleaning and entity/sentiment enrichment themselves.
News is a firehose that's messy to consume programmatically — boilerplate, duplicates, paywalls, inconsistent formats. Our news datasets deliver it clean: article text, headlines, entities and sentiment, deduplicated and structured for NLP and analysis.
License a snapshot or subscribe for a real-time feed. For continuous, filtered news pipelines, see news & media feeds.
One vertical, every platform that matters in it — matched into a single feed your team actually uses.
Real sample structure from this feed. Your free 48-hour sample comes in your category, in this shape — CSV, JSON or straight to your warehouse.
| published_at | source | headline_snippet | entities | sentiment |
|---|---|---|---|---|
| 2026-07-08 | Reuters | Retailer posts record quarter as... | ExampleCo; India | +0.5 |
| 2026-07-08 | Bloomberg | Sector faces margin squeeze amid... | Retail | -0.3 |
| 2026-07-07 | ET | Startup raises Series B for... | StartupX | +0.6 |
| 2026-07-07 | Mint | Regulator proposes new rules on... | Policy | -0.1 |
[
{
"published_at": "2026-07-08",
"source": "Reuters",
"headline_snippet": "Retailer posts record quarter as...",
"entities": "ExampleCo; India",
"sentiment": "+0.5"
},
{
"published_at": "2026-07-08",
"source": "Bloomberg",
"headline_snippet": "Sector faces margin squeeze amid...",
"entities": "Retail",
"sentiment": "-0.3"
},
{
"published_at": "2026-07-07",
"source": "ET",
"headline_snippet": "Startup raises Series B for...",
"entities": "StartupX",
"sentiment": "+0.6"
},
{
"published_at": "2026-07-07",
"source": "Mint",
"headline_snippet": "Regulator proposes new rules on...",
"entities": "Policy",
"sentiment": "-0.1"
}
]
Clean, entity-tagged news text for NLP models and pipelines.
Entity-linked news sentiment on companies and sectors over time.
Structured news data to analyze topics, sources and framing.
Headlines, clean article text (boilerplate removed), publication dates, sources, extracted entities, sentiment scores, category/topic, public author, URL, language, location mentions and timestamps — deduplicated and enriched.
Yes — boilerplate, navigation and ads are stripped, leaving clean article text, and near-duplicates (syndication, mirrors) are removed. Entities and sentiment are added, so the data is analysis-ready rather than raw HTML.
Entity-linked news sentiment on specific companies and sectors, tracked over time, is a signal investors combine with other alt-data. We map mentions to entities and score sentiment so the trend is measurable.
We provide structured metadata, entities, sentiment and excerpts for analysis rather than republishing full copyrighted articles. Licensing terms are set to your use case (NLP training, internal analysis) and agreed up front.