Hugging Face Papers Scraper avatar

Hugging Face Papers Scraper

Pricing

from $9.00 / 1,000 results

Go to Apify Store
Hugging Face Papers Scraper

Hugging Face Papers Scraper

Scrape AI and machine learning research papers from Hugging Face Papers. Get titles, abstracts, authors with affiliations, upvotes, publication dates, ArXiv IDs, and community discussion counts. Search by keyword or browse daily papers.

Pricing

from $9.00 / 1,000 results

Rating

0.0

(0)

Developer

ParseForge

ParseForge

Maintained by Community

Actor stats

1

Bookmarked

7

Total users

2

Monthly active users

a day ago

Last modified

Share

ParseForge Banner

๐Ÿ“„ Hugging Face Papers Scraper

๐Ÿš€ Scrape trending and keyword-searched AI/ML papers from Hugging Face with titles, abstracts, authors, upvotes, arXiv IDs, and GitHub repos. Returns structured data in seconds.

Every day, Hugging Face Papers surfaces the most discussed machine learning research with community upvotes, author profiles, and links to code repositories. This Actor pulls that curated feed or runs keyword searches across the entire index, returning structured records with titles, abstracts, arXiv identifiers, author details, GitHub links, project pages, AI-generated keywords, and community engagement metrics.

Whether you run an AI newsletter, track a research subfield for your lab, or want to spot emerging trends before they go mainstream, this scraper saves you hours of manual browsing. Set it on a daily schedule and let it build a living archive of the papers that matter to your work.

TargetHugging Face Papers
Use CasesResearch newsletters, literature reviews, ML trend tracking, academic monitoring

๐Ÿ“‹ What it does

  • ๐Ÿ“š Paper metadata. Titles, abstracts, arXiv IDs, publication dates, and direct Hugging Face URLs for every paper.
  • ๐Ÿ‘ฅ Author details. Full author lists with Hugging Face usernames and verification status included.
  • โญ Community engagement. Upvote counts, comment totals, and thumbnails so you can gauge which papers resonate.
  • ๐Ÿ’ป Code and project links. GitHub repository URLs and project pages when authors have linked them.
  • ๐Ÿ” Two collection modes. Search by keyword across indexed papers or grab today's trending daily feed.

Each record includes the arXiv ID, paper title, abstract, publication date, full author list with HF handles, upvote and comment counts, thumbnail image, GitHub repo link, project page, and AI-generated keywords.

๐Ÿ’ก Why it matters: Manually checking Hugging Face Papers every day and copying metadata into a spreadsheet takes 30+ minutes. This Actor does it in seconds and delivers a clean, structured dataset ready for analysis.

๐Ÿ“Š Data fields

Each record includes: arxivId, firstAuthor, numAuthors, numComments, publishedAt, scrapedAt, summary, title, upvotes, url. These field names come straight from the actor's dataset schema, so what you see here is what lands in your dataset.

โš ๏ธ Good to Know: Hugging Face Papers indexes new publications daily. Trending mode returns papers curated by the HF team and community for the current day. Search mode queries across all indexed papers. Results are limited by what the Hugging Face API exposes.

๐Ÿš€ How to use

  1. ๐Ÿ“ Sign up. Create a free account with $5 credit (takes 2 minutes).
  2. ๐ŸŒ Open the Actor. Go to the Hugging Face Papers Scraper page on the Apify Store.
  3. ๐ŸŽฏ Set input. Choose a keyword and mode (search or trending), then set your max items.
  4. ๐Ÿš€ Run it. Click Start and let the Actor collect your data.
  5. ๐Ÿ“ฅ Download. Grab your results in the Dataset tab as CSV, Excel, JSON, or XML.

โฑ๏ธ Total time from signup to downloaded dataset: 3-5 minutes. No coding required.

๐Ÿ’ก Pro Tip: browse the complete ParseForge collection for more data scrapers and tools.

โš ๏ธ Disclaimer: this Actor is an independent tool and is not affiliated with, endorsed by, or sponsored by Hugging Face or arXiv. All trademarks mentioned are the property of their respective owners. Only publicly available data is collected.

๐Ÿ†˜ Need Help?

If you hit a bug, have questions about setup, or need a scraper we haven't built yet, open our contact form or write to parseforge@protonmail.com. We also take on paid custom data projects.

For faster answers, join our Discord. It's the best place to get support and suggest new actors.