# US Energy.gov Data Scraper (`parseforge/energy-gov-scraper`) Actor

Scrape energy-related content from Energy.gov, including articles, press releases, documents, titles, dates, offices, and types. Automate collection of structured data from the U.S. Department of Energy, ideal for researchers, journalists, and professionals needing accurate, up-to-date information.

- **URL**: https://apify.com/parseforge/energy-gov-scraper.md
- **Developed by:** [ParseForge](https://apify.com/parseforge) (community)
- **Categories:** Other, Automation, News
- **Stats:** 3 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

Pay per event

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/platform/actors/running/actors-in-store#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

![ParseForge Banner](https://github.com/ParseForge/apify-assets/blob/ad35ccc13ddd068b9d6cba33f323962e39aed5b2/banner.jpg?raw=true)

## 🔬 Energy.gov Scraper

> 🚀 **Collect U.S. Department of Energy articles, press releases, and documents in seconds.** Filter by keyword, office, article type, and language. No coding, no API keys required.

<table><tr>
<td style="border-left:4px solid #0F766E;padding:12px 16px;font-weight:600">Pull structured records from US Energy.gov Data — clean fields ready as CSV, JSON, JSONL, Excel, or XML for downstream pipelines.</td>
</tr></table>

##### Copy to your AI assistant

Copy this block into ChatGPT, Claude, Cursor, or any LLM to start using this actor.

```
parseforge/energy-gov-scraper on Apify. Call: ApifyClient("TOKEN").actor("parseforge/energy-gov-scraper").call(run_input={...}), then client.dataset(run["defaultDatasetId"]).list_items().items for results. Key inputs: maxItems (integer, default 10), keywords (string), articleType (string), language (string), office (string), sort (string, default "date"). Full actor spec: fetch build via GET https://api.apify.com/v2/acts/parseforge~energy-gov-scraper (Bearer TOKEN). Get token: https://console.apify.com/account/integrations
```

The Energy.gov Scraper automates the collection of official content from the U.S. Department of Energy website. It pulls articles, press releases, congressional testimonies, blog posts, success stories, and multimedia content directly from the DOE search system. Each record includes the headline, publication date, source office, content category, direct link, and unique identifier. You can filter by **keyword**, **article type**, **DOE office**, and **language** to zero in on exactly the content you need. Free users can collect up to **10 items** per run, while paid users can retrieve up to **1,000,000 results**.

Whether you are tracking renewable energy policy shifts, monitoring nuclear research announcements, or building a dataset of DOE press releases for media analysis, this tool replaces hours of manual browsing with a single automated run. Results export to **JSON, CSV, or Excel** for immediate use in spreadsheets, dashboards, or data pipelines. Schedule recurring runs to stay current with the latest DOE publications without lifting a finger. The scraper handles pagination, deduplication, and rate limiting automatically so you can focus on analysis instead of data collection.

| Target Audience | Use Cases |
|---|---|
| Policy Analysts | Monitor federal energy policy announcements and congressional testimonies |
| Academic Researchers | Build literature databases from DOE research publications |
| Energy Industry Professionals | Track regulatory changes and press releases by office |
| Journalists | Follow DOE news across topics like renewables, nuclear, and fossil fuels |
| Data Analysts | Export structured DOE content for trend analysis and reporting |
| Government Affairs Teams | Stay current on DOE initiatives and funding announcements |

### 📋 What the Energy.gov Scraper does

- 📝 **Article headlines** - capture the title of every article, press release, blog post, or document published on energy.gov
- 🔗 **Direct URLs** - collect working links to each piece of content for quick reference or archival
- 📅 **Publication dates** - track when content was published to build timelines and spot trends
- 👤 **Source offices** - identify which DOE office or organization published the content (e.g., Office of Energy Efficiency and Renewable Energy)
- 🎯 **Content categories** - classify each item by type: blog, press release, document, success story, congressional testimony, or multimedia
- 🆔 **Unique identifiers** - get UUIDs for each article to manage deduplication and data integrity

The scraper connects to the DOE search system and iterates through results using your specified filters. It collects structured data from each listing, normalizes timestamps, and removes duplicate entries using unique article IDs. All results are pushed to an Apify dataset in real time, so you can preview data as the run progresses.

> 💡 **Why it matters:** Energy.gov publishes thousands of articles annually across dozens of offices. Manually tracking this content is impractical. This scraper gives you structured, filterable access to the entire catalog in minutes.

### 📊 Data fields

Each record includes: `date`, `offices`, `scrapedTimestamp`, `title`, `type`, `url`, `uuid`. All 7 field names come from a real production run, so what you see here is what lands in your dataset.

> ⚠️ **Good to Know:** Free users are automatically limited to 10 items per run. Leave `keywords` empty to browse all available content. The `articleType` field uses numeric codes internally, but you can also use descriptive names.

### 🚀 How to use

1. **Sign up** - [Create a free Apify account with $5 credit](https://console.apify.com/sign-up?fpr=vmoqkp)
2. **Find the Actor** - Search for "Energy.gov Scraper" in the Apify Store
3. **Configure your filters** - Set keywords, article type, office, language, and max items
4. **Start the run** - Click "Start" and watch results appear in real time
5. **Export your data** - Download as JSON, CSV, or Excel from the dataset tab

> 🕒 **Typical run time:** 30 seconds to 2 minutes for up to 100 items. Larger runs with 500+ items may take 5 to 10 minutes.

### 🔗 Recommended Actors

| Actor | Description |
|---|---|
| [USAspending Scraper](https://apify.com/parseforge/usaspending-scraper) | Extract federal spending data and contract information from USAspending.gov |
| [GSA eLibrary Scraper](https://apify.com/parseforge/gsa-elibrary-scraper) | Collect government contractor and vendor data from the GSA eLibrary |
| [PR Newswire Scraper](https://apify.com/parseforge/pr-newswire-scraper) | Collect press releases and news articles from PR Newswire |
| [FINRA BrokerCheck Scraper](https://apify.com/parseforge/finra-brokercheck-scraper) | Search broker and firm registration data from the FINRA registry |
| [FAA Aircraft Registry Scraper](https://apify.com/parseforge/faa-aircraft-registry-scraper) | Look up aircraft registration records by N-number from the FAA |

> 💡 **Pro Tip:** Combine the Energy.gov Scraper with the USAspending Scraper to cross-reference DOE announcements with actual federal spending data.

> **Disclaimer:** This Actor is an independent tool and is not affiliated with, endorsed by, or sponsored by the U.S. Department of Energy or Energy.gov. All trademarks mentioned are the property of their respective owners.

### 🆘 Need Help?

If you hit a bug, have questions about setup, or need a scraper we haven't built yet, open our [contact form](https://tally.so/r/BzdKgA) or write to parseforge@protonmail.com. We also take on paid custom data projects.

For faster answers, join our [Discord](https://parseforge.co/discord). It's the best place to get support and suggest new actors.

# Actor input Schema

## `maxItems` (type: `integer`):

Maximum number of articles to collect per run.

## `keywords` (type: `string`):

Search keywords to filter articles.

## `articleType` (type: `string`):

Filter by article type

## `language` (type: `string`):

Filter by language

## `office` (type: `string`):

Filter by DOE office or organization

## `sort` (type: `string`):

Sort order for results

## Actor input object example

```json
{
  "maxItems": 10,
  "sort": "date"
}
```

# Actor output Schema

## `articles` (type: `string`):

Complete dataset with all scraped articles including titles, URLs, dates, offices, and article types

## `overview` (type: `string`):

Overview view of articles with key fields displayed in a table format

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "maxItems": 10
};

// Run the Actor and wait for it to finish
const run = await client.actor("parseforge/energy-gov-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "maxItems": 10 }

# Run the Actor and wait for it to finish
run = client.actor("parseforge/energy-gov-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print("💾 Check your data here: https://console.apify.com/storage/datasets/" + run["defaultDatasetId"])
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "maxItems": 10
}' |
apify call parseforge/energy-gov-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "command": "npx",
            "args": [
                "mcp-remote",
                "https://mcp.apify.com/?tools=parseforge/energy-gov-scraper",
                "--header",
                "Authorization: Bearer <YOUR_API_TOKEN>"
            ]
        }
    }
}

```

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/acts/ltZCMgiYxFhFJhE8N/builds/fx0noFnKg1BuJxfnA/openapi.json
