# Reddit Scraper - Posts, Comments, Subreddits & User Profiles (`convertfleetdotonline/reddit-scraper`) Actor

Reddit scraper for posts, comments, subreddit communities and user profiles by keyword or URL. NSFW filter, sort and time-range options. Fast, scalable Reddit data extraction to JSON, CSV or Excel.

- **URL**: https://apify.com/convertfleetdotonline/reddit-scraper.md
- **Developed by:** [Hasnain Nisar](https://apify.com/convertfleetdotonline) (community)
- **Categories:** Social media, Lead generation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$1.50 / 1,000 item scrapeds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/platform/actors/running/actors-in-store#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## Reddit Scraper - Scrape Reddit Posts, Comments, Subreddits & Users

**Scrape Reddit at scale with this Reddit scraper — a web scraper for Reddit posts, comments, subreddit communities, and user profiles, run entirely via the Apify API.** Pull data by keyword or URL — with NSFW filtering, sort order, and time-range controls. It is a simple way to scrape Reddit without wrestling with the official Reddit API.

### What this Reddit Scraper does

The actor uses Reddit's public JSON API (the same one that powers `old.reddit.com` and is freely accessible) and supports four search dimensions in one run:

- **Posts** — keyword search across all of Reddit or restricted to a subreddit
- **Comments** — keyword search inside Reddit comment bodies
- **Communities (subreddits)** — find subreddits matching a keyword
- **URLs** — pass any Reddit URL (subreddit, single post, user profile) and the actor routes it correctly

You also get optional **comment scraping per post**, NSFW filtering, sort by new/hot/top/relevance/comments, and time-range filters (hour/day/week/month/year/all).

### Why use this Reddit Scraper?

- **Runs fully via the Apify API** — bypass OAuth and the new pricing tier
- **No rate-limit headaches** — the actor handles 429 backoff automatically
- **One-shot multi-mode** — search posts + comments + communities in a single run
- **URL-aware** — paste subreddit / post / user URLs and let the actor route
- **NSFW toggle** — opt in or out cleanly
- **Comment expansion** — pull top-level comments per post in the same run

### Use cases

- **Brand monitoring** — track every mention of your brand or product
- **Sentiment analysis** — feed Reddit comments into your AI/ML pipeline
- **Audience research** — find which subreddits your audience hangs out in
- **Trend tracking** — surface trending posts in your niche
- **Influencer discovery** — find prolific posters in target subreddits
- **Content ideation** — see what's getting upvoted in your category
- **Lead generation** — spot people asking for the product you sell
- **Academic research** — public-domain data for social science studies

### Input

Search by keyword:

```json
{
    "keywords": ["bitcoin", "ethereum"],
    "searchPosts": true,
    "searchComments": false,
    "searchCommunities": true,
    "communityFilter": "",
    "sortBy": "new",
    "timeRange": "week",
    "includeNsfw": false,
    "maxPosts": 100,
    "maxCommunities": 10
}
```

Scrape a specific subreddit:

```json
{
    "urls": ["https://www.reddit.com/r/programming"],
    "scrapeComments": true,
    "maxPosts": 50,
    "maxCommentsPerPost": 20
}
```

Scrape a specific post + its comments:

```json
{
    "urls": ["https://www.reddit.com/r/programming/comments/abc123/"],
    "scrapeComments": true
}
```

### Output

Each post:

```json
{
    "type": "post",
    "id": "abc123",
    "title": "Title of the post",
    "body": "Self-text (empty for link posts)",
    "subreddit": "programming",
    "author": "username",
    "ups": 4521,
    "num_comments": 234,
    "created_utc": 1735689600,
    "url": "https://www.reddit.com/r/programming/comments/abc123/title/",
    "is_nsfw": false,
    "score": 4521,
    "link_flair": "Discussion"
}
```

Each comment has `type: "comment"`. Each community has `type: "community"` with `num_comments` repurposed as `subscribers`.

### How it works

The actor uses Reddit's public JSON API (`reddit.com/...json` endpoints) — the same machine-readable surface Reddit has supported since 2009. No login, no OAuth, no developer account. Built-in 429 retry-after handling keeps runs reliable.

### Cost & speed

A 100-post keyword search completes in 5–10 seconds. With comment expansion (10 posts × 10 comments each) it takes 20–30 seconds. Memory usage stays under 256 MB.

### Related actors

- **Twitter / X Scraper** — bulk tweet scraping
- **TikTok Scraper** — TikTok by hashtag, profile, search
- **LinkedIn Profile Scraper** — bulk LinkedIn profile extraction
- **YouTube Channel Scraper** — bulk YouTube video metadata

### FAQ

**Q: How do I scrape Reddit posts and comments?**
Search by keyword with `searchPosts` and `searchComments`, or pass Reddit URLs. Enable `scrapeComments` (with optional `maxCommentsPerPost`) to pull top-level comments for each post in the same run. Posts, comments, and communities can all be searched in a single job.

**Q: Can I scrape a whole subreddit or a user profile?**
Yes. Pass any Reddit URL — a subreddit, a single post, or a user profile — into `urls` and the actor routes it correctly. You can also target a subreddit by keyword using `communityFilter`, with sort and time-range controls.

**Q: Do I need the Reddit API?**
No. The actor uses Reddit's public JSON endpoints (`reddit.com/...json`) with automatic 429 back-off — no OAuth, developer account, or paid API tier required.

**Q: Do I need a Reddit account?**
No.

**Q: Will Reddit ban me?**
No — the actor uses public JSON endpoints with realistic browser headers and respects `Retry-After` on 429 responses.

**Q: Can I scrape entire subreddits historically?**
The Reddit listing API caps at ~1,000 posts per sort order. For deeper history, run the actor multiple times with different `sortBy` values (new, top, hot) to capture different slices.

**Q: How accurate is the comment scrape?**
Top-level comments only — Reddit returns these in the post's `.json` endpoint. Nested replies are not currently expanded.

**Q: Is this legal?**
Reddit's public JSON endpoints are explicitly accessible to anyone. Always comply with Reddit's User Agreement and your local data-protection laws when storing personal data.

# Actor input Schema

## `keywords` (type: `array`):

Search terms to query Reddit for. Leave empty if scraping by URL only.

## `urls` (type: `array`):

Direct Reddit URLs - subreddits (/r/<name>), single posts (/comments/<id>), or user profiles (/user/<name>).

## `searchPosts` (type: `boolean`):

Include posts in keyword searches.

## `searchComments` (type: `boolean`):

Include comments in keyword searches.

## `searchCommunities` (type: `boolean`):

Include subreddit communities in keyword searches.

## `communityFilter` (type: `string`):

Optional - restrict keyword searches to a single subreddit (without /r/).

## `sortBy` (type: `string`):

Reddit sort order applied to keyword searches.

## `timeRange` (type: `string`):

Time window for keyword searches (only applies to certain sort orders).

## `scrapeComments` (type: `boolean`):

If true, also pull top-level comments for each post returned.

## `includeNsfw` (type: `boolean`):

Include results flagged as NSFW (over\_18).

## `maxPosts` (type: `integer`):

Maximum number of posts to return per query/URL.

## `maxComments` (type: `integer`):

Maximum number of comments to return when scraping comment search.

## `maxCommentsPerPost` (type: `integer`):

When 'Scrape comments per post' is on, the max comments returned per post.

## `maxCommunities` (type: `integer`):

Maximum number of subreddit communities to return per community search.

## `cookies` (type: `string`):

Reddit returns <b>403 Blocked</b> for anonymous requests coming from datacenter IPs, which is what Apify's default proxy uses. Pasting the cookies of a signed-in Reddit session removes that block without needing residential proxies.<br><br><b>How to get them:</b> sign in to reddit.com → open DevTools (F12) → <b>Network</b> tab → reload → click the first request → copy the whole <code>cookie:</code> request header and paste it here. A JSON array from a cookie-exporter extension also works.<br><br>Leave empty to try anonymously (works with residential proxies).

## `proxyConfiguration` (type: `object`):

Routes Reddit requests through Apify Proxy. Reddit blocks datacenter IPs — residential proxy is required.

## Actor input object example

```json
{
  "keywords": [
    "bitcoin"
  ],
  "urls": [],
  "searchPosts": true,
  "searchComments": false,
  "searchCommunities": false,
  "communityFilter": "",
  "sortBy": "new",
  "timeRange": "all",
  "scrapeComments": false,
  "includeNsfw": false,
  "maxPosts": 25,
  "maxComments": 10,
  "maxCommentsPerPost": 10,
  "maxCommunities": 5,
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ]
  }
}
```

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "keywords": [
        "bitcoin"
    ],
    "proxyConfiguration": {
        "useApifyProxy": true,
        "apifyProxyGroups": [
            "RESIDENTIAL"
        ]
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("convertfleetdotonline/reddit-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "keywords": ["bitcoin"],
    "proxyConfiguration": {
        "useApifyProxy": True,
        "apifyProxyGroups": ["RESIDENTIAL"],
    },
}

# Run the Actor and wait for it to finish
run = client.actor("convertfleetdotonline/reddit-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print("💾 Check your data here: https://console.apify.com/storage/datasets/" + run["defaultDatasetId"])
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "keywords": [
    "bitcoin"
  ],
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ]
  }
}' |
apify call convertfleetdotonline/reddit-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "command": "npx",
            "args": [
                "mcp-remote",
                "https://mcp.apify.com/?tools=convertfleetdotonline/reddit-scraper",
                "--header",
                "Authorization: Bearer <YOUR_API_TOKEN>"
            ]
        }
    }
}

```

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/acts/yWiLk6pZyo8T5CUK4/builds/PO2VCZt6a3KbogrTB/openapi.json
