# Reddit Post Comments Scraper | $2.5/1K Comments (`ethereal_wool/reddit-post-comments`) Actor

Extract Reddit post comments data — title, author, engagement, and more. Scrape by keyword, URL or ID. Export to JSON, CSV & Excel, use the API, schedule runs and integrate. No code required.

- **URL**: https://apify.com/ethereal\_wool/reddit-post-comments.md
- **Developed by:** [Jackie Chen](https://apify.com/ethereal_wool) (community)
- **Categories:** Social media, Marketing
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$3.00 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/platform/actors/running/actors-in-store#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## Reddit Post Comments Scraper

![reddit-post-comments](https://api.apify.com/v2/key-value-stores/48FFrfZnw5XqzUC1C/records/reddit-post-comments__hero)

Scrape the **full comment forest of any Reddit post**. Give one or more post IDs and the
Actor returns one clean, structured item per comment: the comment text (markdown), score,
author, depth, parent ID, timestamps, and removed / locked / stickied flags. It can
optionally follow "load more replies" to expand deep nested threads and attach the post's
metadata to every comment.

> **Unofficial.** This Actor is not affiliated with, authorized, or endorsed by Reddit, Inc.
> It is an independent tool that retrieves publicly available data via a third-party API.
> Use it in compliance with Reddit's terms and all applicable laws; you are responsible for
> how you use the retrieved data.

### What it does

- **Comment forest** — for each post, fetches the comment tree sorted by Best / Top / New /
  Controversial / Old / Q\&A.
- **Nested replies (optional)** — follows each thread's "load more replies" cursor to pull
  deeper comments that aren't in the initial response.
- **Post metadata (optional)** — fetches the post's title, subreddit, author, score and
  comment count once and attaches a `post` object to every comment item.

### Input

| Field | Type | Default | Description |
|---|---|---|---|
| `postIds` | string\[] | `["t3_bawfcs"]` | Reddit post fullnames (`t3_…`). A bare id or a full post URL is also accepted. |
| `sortType` | enum | `CONFIDENCE` | `CONFIDENCE` (Best) / `TOP` / `NEW` / `CONTROVERSIAL` / `OLD` / `QA`. |
| `maxItems` | integer | `50` | Max total comments across all posts. |
| `expandReplies` | boolean | `false` | Follow "load more" cursors to fetch deeper nested replies. |
| `includePostInfo` | boolean | `false` | Attach post metadata (title, subreddit, author, score) to each comment. |

#### Example input

```json
{
  "postIds": ["t3_bawfcs", "https://www.reddit.com/r/AskReddit/comments/abc123/some_thread/"],
  "sortType": "TOP",
  "maxItems": 500,
  "expandReplies": true,
  "includePostInfo": true
}
```

### Output

One dataset item per comment:

```json
{
  "id": "t1_ekf5nop",
  "postId": "t3_bawfcs",
  "parentId": null,
  "depth": 0,
  "childCount": 12,
  "content": "There was one I saw a while ago where it was ...",
  "score": 3633,
  "author": "Cobalt-Royal",
  "authorId": "t2_pr5g8",
  "createdAt": "2019-04-08T21:29:42.750000+0000",
  "editedAt": "2019-04-09T00:06:28.303000+0000",
  "isRemoved": false,
  "isLocked": false,
  "isStickied": false,
  "isArchived": true,
  "distinguishedAs": null,
  "permalink": "/r/AskReddit/comments/bawfcs/.../ekf5nop/",
  "url": "https://www.reddit.com/r/AskReddit/comments/bawfcs/.../ekf5nop/",
  "hasMoreReplies": true,
  "post": {
    "id": "t3_bawfcs",
    "title": "What's the creepiest Ask Reddit thread you have come across?",
    "subreddit": "r/AskReddit",
    "author": null,
    "score": null,
    "commentCount": 1234
  }
}
```

(`post` is present only when `includePostInfo` is enabled.)

### Notes

- Comment IDs are de-duplicated within a run.
- Data is sourced live; Reddit's edge occasionally rate-limits, so the Actor retries
  transient blocks with exponential backoff.
- `post_id` must reference the `t3_` post fullname; the Actor normalizes bare ids and URLs
  for you.

### Quick start

1. Open the Actor and press **Run** — the default input works out of the box.
2. Adjust the input fields below to your target (keywords, IDs, or URLs) and set `maxItems` to cap spend.
3. Grab results from the **Dataset** tab as JSON / CSV / Excel, or pull them via the [Apify API](https://docs.apify.com/api/v2) and MCP from your own code.

No proxies to configure, no cookies to paste, no login — the Actor handles everything server-side.

### Why teams switch to this Reddit comments scraper

Comment trees are where Reddit's real signal lives, but they're the part most
scrapers handle worst — browser-based actors time out on big threads and
charge $15–20 per 1,000 comments. This Actor walks the thread via a direct
HTTP API and returns the tree as flat, parent-linked JSON at **$2.50 per
1,000 comments**.

### What people build with it

- **Sentiment analysis** — feed a product-launch thread's comments to an LLM
  and get a structured verdict on how it landed, weighted by comment score.
- **Voice-of-customer mining** — the complaints and praise under reviews and
  comparison threads are unfiltered product feedback you can't survey for.
- **Conversation datasets** — parent-child comment pairs are natural dialogue
  data for fine-tuning chat models.
- **Crisis monitoring** — when a thread about your brand takes off, pull the
  tree hourly and track which concerns are gaining score.
- **Summarized digests** — schedule the Actor + an LLM step to turn daily
  megathreads into a morning brief.

### Tips for better results

- Accept either the full post URL or the bare post ID.
- Each comment carries `parentId`, `depth`, and `score`, so you can rebuild
  the tree, filter to top-level only, or weight by votes.
- Find the threads worth expanding with
  [Reddit Post Search](https://apify.com/ethereal_wool/reddit-post-search) or
  [Reddit Subreddit Posts](https://apify.com/ethereal_wool/reddit-subreddit-posts).

***

### Why this Actor

- **Direct API, no headless browser** — fast, stable runs with nothing to babysit.
- **No login, no cookies** — we never touch your accounts, so there's no ban risk.
- **Fresh, real-time data** — every run reads the source live, not a stale cache.
- **Pay per result** — you're billed only for the rows actually delivered.
- **Structured JSON** — export to CSV, Excel, or JSON, or pull straight from the API / MCP.

### Use cases

- Mine audience sentiment and feature requests from real comment threads.
- Surface the most-liked replies and frequent questions under any post.
- Build moderation, UGC, or social-listening datasets at scale.
- Spot superfans and detractors by author and engagement.

### FAQ

**Do I need an account, cookies, or to log in anywhere?**
No. The Actor talks to a fast, direct HTTP API server-side — you just provide inputs and run it.

**How am I billed?**
Pay-per-result: a fixed price per row returned, with no separate platform/compute charge. Caps like `maxItems` keep spend predictable.

**Can I run it on a schedule or call it from my app?**
Yes — use Apify Schedules, the REST API, the JavaScript / Python clients, or the MCP server. See the **API** tab.

**Is this affiliated with Reddit?**
No. It's an independent tool that collects publicly available data. Use it in line with the platform's terms and applicable law.

### More Reddit scrapers by us

- [**Reddit Post Search**](https://apify.com/ethereal_wool/reddit-post-search) — Keyword post search across Reddit
- [**Reddit Subreddit Posts**](https://apify.com/ethereal_wool/reddit-subreddit-posts) — Posts from any subreddit · sort & time
- [**Reddit User Posts**](https://apify.com/ethereal_wool/reddit-user-posts) — A user's posts & comments history

Browse the full fleet → **https://apify.com/ethereal\_wool**

# Actor input Schema

## `postIds` (type: `array`):

Reddit post fullnames to scrape comments from, e.g. "t3\_bawfcs". A bare id ("bawfcs") or a full post URL is also accepted. Each post's comment forest is fetched and paginated independently.

## `sortType` (type: `string`):

How to sort the comment forest.

## `maxItems` (type: `integer`):

Maximum total number of comments to scrape across all posts.

## `expandReplies` (type: `boolean`):

Follow 'load more replies' cursors to fetch deeper nested replies that aren't in the initial comment forest. Slower and uses more API calls.

## `includePostInfo` (type: `boolean`):

Fetch each post's metadata (title, subreddit, author, score, comment count) once and attach a 'post' object to every comment item.

## `proxyConfiguration` (type: `object`):

Optional. Route the upstream API calls through an Apify Proxy to vary the source IP. Usually not needed.

## Actor input object example

```json
{
  "postIds": [
    "t3_bawfcs"
  ],
  "sortType": "CONFIDENCE",
  "maxItems": 50,
  "expandReplies": false,
  "includePostInfo": false,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "postIds": [
        "t3_bawfcs"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("ethereal_wool/reddit-post-comments").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "postIds": ["t3_bawfcs"] }

# Run the Actor and wait for it to finish
run = client.actor("ethereal_wool/reddit-post-comments").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print("💾 Check your data here: https://console.apify.com/storage/datasets/" + run["defaultDatasetId"])
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "postIds": [
    "t3_bawfcs"
  ]
}' |
apify call ethereal_wool/reddit-post-comments --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "command": "npx",
            "args": [
                "mcp-remote",
                "https://mcp.apify.com/?tools=ethereal_wool/reddit-post-comments",
                "--header",
                "Authorization: Bearer <YOUR_API_TOKEN>"
            ]
        }
    }
}

```

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/acts/nctmTAZcsnEWbEhSC/builds/KP0bzpf6TT6O5TSEV/openapi.json
