# PubMed Search Scraper (`easyapi/pubmed-search-scraper`) Actor

Scrape research papers and academic articles from PubMed based on search terms. Extract comprehensive article metadata including titles, authors, citations, abstracts, and more. Perfect for medical research and literature reviews.

- **URL**: https://apify.com/easyapi/pubmed-search-scraper.md
- **Developed by:** [EasyApi](https://apify.com/easyapi) (community)
- **Categories:** Integrations
- **Stats:** 82 total users, 7 monthly users, 100.0% runs succeeded, 6 bookmarks
- **User rating**: 5.00 out of 5 stars

## Pricing

from $2.99 / 1,000 results

This Actor is paid per event and usage. You are charged both the fixed price for specific events and for Apify platform usage.

Learn more: https://docs.apify.com/platform/actors/running/actors-in-store#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## PubMed Search Scraper 🔬

### 📋 Overview

Extract academic articles and research papers from PubMed, the world's leading database of biomedical literature. This actor allows you to scrape detailed information from search results based on your keywords.

### ✨ Features

- 🔎 Scrape articles based on custom search queries
- 📑 Extract comprehensive article metadata:
  - Title and article ID
  - Full & short author lists
  - Complete & abbreviated journal citations
  - PMID (PubMed Identifier)
  - Article tags and types
  - Full & truncated abstracts
  - Social sharing links
- ⚡ High-performance scrolling pagination
- 🛡️ Built-in anti-blocking measures
- 🎯 Configurable maximum items limit

### 💡 Use Cases

- Medical research and literature reviews
- Academic meta-analyses
- Tracking research trends
- Building research databases
- Bibliometric analysis
- Scientific data mining

### 📤 Output

The actor outputs detailed article information in JSON format, including:

- Article title and unique identifier
- Author information (full and short formats)
- Journal citation details
- PMID reference
- Article type tags
- Abstract content
- Social sharing links

### 💪 Tips for Optimal Usage

1. Use specific search terms for more targeted results
2. Consider breaking large searches into smaller queries
3. Allow sufficient run time for larger result sets
4. Monitor your usage to stay within PubMed's guidelines

### 🔗 Links

- [PubMed Home](https://pubmed.ncbi.nlm.nih.gov/)
- [Search Syntax Guide](https://pubmed.ncbi.nlm.nih.gov/help/#search-tags)

#### Input Example

A full explanation of an input example in JSON.

```
{
    "searchUrls": ["https://pubmed.ncbi.nlm.nih.gov/?term=rheumatoid%20arthritis"],
    "maxItems": 30
}
```

#### Output sample

The results will be wrapped into a dataset which you can always find in the **Storage** tab. Here's an excerpt from the data you'd get if you apply the input parameters above:

And here is the same data but in JSON. You can choose in which format to download your data: JSON, JSONL, Excel spreadsheet, HTML table, CSV, or XML.

```
[
	{
		"title": "Rheumatoid arthritis.",
		"articleId": "27156434",
		"articleUrl": "https://pubmed.ncbi.nlm.nih.gov/27156434/",
		"authors": {
			"full": "Smolen JS, Aletaha D, McInnes IB.",
			"short": "Smolen JS, et al."
		},
		"citation": {
			"full": "Lancet. 2016 Oct 22;388(10055):2023-2038. doi: 10.1016/S0140-6736(16)30173-8. Epub 2016 May 3.",
			"short": "Lancet. 2016."
		},
		"pmid": "27156434",
		"tags": [
			"Free article.",
			"Review."
		],
		"abstract": {
			"full": "Rheumatoid arthritis is a chronic inflammatory joint disease, which can cause cartilage and bone damage as well as disability. ...In this Seminar, we describe current insights into genetics and aetiology, pathophysiology, epidemiology, assessment, therapeutic agents …",
			"short": "Rheumatoid arthritis is a chronic inflammatory joint disease, which can cause cartilage and bone damage as well as disability. …"
		},
		"shareLinks": {
			"twitter": "http://twitter.com/intent/tweet?text=Rheumatoid%20arthritis.%20https%3A//pubmed.ncbi.nlm.nih.gov/27156434/",
			"facebook": "http://www.facebook.com/sharer/sharer.php?u=https%3A//pubmed.ncbi.nlm.nih.gov/27156434/",
			"permalink": "https://pubmed.ncbi.nlm.nih.gov/27156434/"
		}
	},
	{
		"title": "Management of Rheumatoid Arthritis: An Overview.",
		"articleId": "34831081",
		"articleUrl": "https://pubmed.ncbi.nlm.nih.gov/34831081/",
		"authors": {
			"full": "Radu AF, Bungau SG.",
			"short": "Radu AF, et al."
		},
		"citation": {
			"full": "Cells. 2021 Oct 23;10(11):2857. doi: 10.3390/cells10112857.",
			"short": "Cells. 2021."
		},
		"pmid": "34831081",
		"tags": [
			"Free PMC article.",
			"Review."
		],
		"abstract": {
			"full": "Rheumatoid arthritis (RA) is a multifactorial autoimmune disease of unknown etiology, primarily affecting the joints, then extra-articular manifestations can occur. ...",
			"short": "Rheumatoid arthritis (RA) is a multifactorial autoimmune disease of unknown etiology, primarily affecting the joints, then ext …"
		},
		"shareLinks": {
			"twitter": "http://twitter.com/intent/tweet?text=Management%20of%20Rheumatoid%20Arthritis%3A%20An%20Overview.%20https%3A//pubmed.ncbi.nlm.nih.gov/34831081/",
			"facebook": "http://www.facebook.com/sharer/sharer.php?u=https%3A//pubmed.ncbi.nlm.nih.gov/34831081/",
			"permalink": "https://pubmed.ncbi.nlm.nih.gov/34831081/"
		}
	},
    ...
]
```

# Actor input Schema

## `searchUrls` (type: `array`):

Array of PubMed search URLs to scrape

## `maxItems` (type: `integer`):

Maximum number of articles to scrape

## Actor input object example

```json
{
  "searchUrls": [
    "https://pubmed.ncbi.nlm.nih.gov/?term=cancer"
  ],
  "maxItems": 20
}
```

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {};

// Run the Actor and wait for it to finish
const run = await client.actor("easyapi/pubmed-search-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {}

# Run the Actor and wait for it to finish
run = client.actor("easyapi/pubmed-search-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print("💾 Check your data here: https://console.apify.com/storage/datasets/" + run["defaultDatasetId"])
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{}' |
apify call easyapi/pubmed-search-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "command": "npx",
            "args": [
                "mcp-remote",
                "https://mcp.apify.com/?tools=easyapi/pubmed-search-scraper",
                "--header",
                "Authorization: Bearer <YOUR_API_TOKEN>"
            ]
        }
    }
}

```

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/acts/Vj1A9bBFIIVF5VH6m/builds/basylC9YeQDQZzJZ2/openapi.json
