Google Ads Transparency Scraper avatar

Google Ads Transparency Scraper

Pricing

Pay per event

Go to Apify Store
Google Ads Transparency Scraper

Google Ads Transparency Scraper

Scrape ad creatives from the Google Ads Transparency Center by advertiser domain or advertiser ID — creative, format, regions, first/last shown, landing URL — export to JSON or CSV. A Google Ads Transparency API alternative and data exporter. You pay only for ads that land.

Pricing

Pay per event

Rating

0.0

(0)

Developer

DevilScrapes

DevilScrapes

Maintained by Community

Actor stats

0

Bookmarked

43

Total users

14

Monthly active users

3 days ago

Last modified

Share


Quick answer: The Google Ads Transparency Scraper pulls every public ad creative Google logs for a brand or advertiser — creative ID, format, landing domain, impression counts, first/last-seen dates, and preview URLs — straight from the Google Ads Transparency Center (the closest thing Google has to Meta's "Ad Library") into JSON, CSV, or Excel. Google ships no official API for this data; we replay its internal RPC ourselves and absorb the fingerprinting, proxy rotation, and retries so you don't have to. Pricing is pay-per-result — $3.01 per 1 000 ads landed, no subscription, no card required to try it.

🎯 What this scrapes

The Google Ads Transparency Center is Google's public registry of every ad campaign running on Search, YouTube, Display, Shopping, Maps, and Play. This Google Ads Transparency Scraper talks directly to Google's internal SearchService/SearchCreatives RPC — so it pulls fast, stable structured data without the overhead of a full browser session.

Google publishes no official API for this data. We reverse-engineered the RPC, replay it with a real-browser TLS fingerprint, and absorb all the reliability work so you get clean rows.

Per creative you get:

FieldTypeNotes
advertiser_idstringGoogle's stable advertiser identifier (e.g. AR0123…)
advertiser_namestringPublic-facing brand name
creative_idstringStable per-creative ID (e.g. CR0123…)
creative_urlstringDeep link into Transparency Center
landing_domainstringClick-through domain
format_typeintegerNumeric format code (1=text, 2=image, 3=video — inferred)
first_shown_tsintegerUnix seconds, first observed impression
last_shown_tsintegerUnix seconds, last observed impression
impressionsintegerGoogle-reported impression count
preview_image_urlstring | nullStatic thumbnail (image creatives)
preview_content_js_urlstring | nullJS bundle URL (video/rich creatives)
regionstringLocale label you passed (display only)
scraped_atstringISO-8601 UTC timestamp

Why we replay the RPC instead of automating a browser

Most listings in this niche point a headless browser at the public Transparency Center UI and re-parse the rendered DOM on every run — one slow page load per handful of creatives, and a broken selector every time Google reshuffles its frontend. We skip the browser: our client speaks the internal SearchCreatives RPC directly, replaying the exact TLS handshake a real Chrome tab would send. One round-trip returns roughly 40 structured creatives, and because it's a backend contract rather than rendered HTML, a Google frontend redesign doesn't break us the way it breaks DOM scrapers overnight.

We're also upfront about what the RPC can't do: Google's SearchCreatives endpoint ignores the region parameter server-side — we tested every plausible request shape to confirm it, and we'd rather publish that finding than ship a region filter that quietly does nothing (see Limitations).

🔥 What we handle for you

  • 🛡️ Browser fingerprint rotation — every RPC call replays a real Chrome / Firefox TLS handshake via curl-cffi, so Google's edge sees a browser, not a Python script.
  • 🌐 Residential proxy rotation via Apify Proxy — sticky sessions carry cookie continuity across hundreds of pages, and we cut over to a fresh session and exit IP the moment one gets rate-limited.
  • 🔁 Retries with exponential backoff on 408 / 429 / 5xx — up to 5 attempts per page, Retry-After honoured.
  • 🧱 Rate-limit-aware pacing — when Google's RPC pushes back, we slow the crawl down instead of burning the session.
  • 🧊 Clean, typed dataset rows — Pydantic-validated schema, golden-file tested against four real creative shapes (static image, rich video, minimal, malformed) before every release.
  • 💰 Pay-per-result pricing — a small per-run fee, then you're charged only for ad rows that actually land in your dataset. No data, no charge.
  • 🧪 Batch input, deduplicated — scrape dozens of domains and advertiser IDs in a single run; overlapping creatives across targets are merged, not double-billed.

💡 Use cases

  • Competitor ad-spend tracking — pull every Nike ad once a week and diff the creative set to see what launched.
  • Trademark enforcement — monitor advertisers running ads against your brand keyword; combine with your own takedown workflow.
  • Affiliate-fraud detection — flag advertisers whose landing domain doesn't match the advertiser name.
  • Political-ad monitoring — track which advertisers are active in an election cycle.
  • Brand-safety audits — for agencies, prove the ads currently live for a client before the QBR.
  • Market research — observe how saturated a vertical (crypto, sports betting, supplements) is with active creatives.
  • AI / RAG ingestion — feed creative metadata and image URLs into a vector store for image-grounded competitive analysis.

⚙️ How to use it

  1. Click "Try for free" at the top of the page.
  2. Paste one or more brand domains into the Brand domains field (e.g. nike.com, adidas.com). One per line. Each domain spawns its own scrape.
  3. (Optional) Drop in advertiser IDs if you already know them — they look like AR0123456789 and live in the Transparency Center URL when you click into an advertiser.
  4. (Optional) Set a date window to narrow ad activity. Defaults to the last 365 days.
  5. Run. Each ad is one row in the dataset; export to JSON, CSV, or Excel from the Storage tab.

The first run on a new account uses $5 of free Apify credit — that's roughly 1 650 ads at our pricing.

📥 Input

The schema lives in .actor/input_schema.json. The fields:

FieldTypeRequiredDefaultNotes
searchDomainsarray of stringone of["nike.com"]Brand landing domains, one per line
advertiserIdsarray of stringone of[]Google advertiser IDs (AR…)
regionenum stringnoanywhereDisplay-only — Google's RPC does not filter by region (see Limitations)
dateFromstring (YYYY-MM-DD)no365 days backLower bound of ad-activity window
dateTostring (YYYY-MM-DD)notoday (UTC)Upper bound
maxResultsintegerno1000Total dataset items across all targets. 0 = unlimited
maxPagesintegerno25RPC budget per target (40 ads × 25 pages = 1 000 / target)
proxyConfigurationproxy confignoApify Proxy enabledSticky session recommended

At least one of searchDomains or advertiserIds must contain at least one entry.

Example input

{
"searchDomains": ["nike.com", "adidas.com"],
"advertiserIds": ["AR03012025048987521025"],
"region": "US",
"dateFrom": "2025-11-15",
"dateTo": "2026-05-15",
"maxResults": 5000,
"maxPages": 25,
"proxyConfiguration": { "useApifyProxy": true }
}

📤 Output

Every row is one creative. Example:

{
"advertiser_id": "AR18378488041124659201",
"advertiser_name": "Nike Retail BV",
"creative_id": "CR15771942603307614209",
"creative_url": "https://adstransparency.google.com/advertiser/AR18378488041124659201/creative/CR15771942603307614209?region=anywhere",
"landing_domain": "nike.com",
"format_type": 1,
"first_shown_ts": 1761145807,
"last_shown_ts": 1778871417,
"impressions": 205,
"preview_image_url": "https://tpc.googlesyndication.com/archive/simgad/12774179880874022668",
"preview_content_js_url": null,
"region": "anywhere",
"scraped_at": "2026-05-15T19:17:59+00:00"
}

Export options once the run finishes:

  • JSON — full payload, ideal for AI/RAG pipelines
  • CSV / Excel — for analyst spreadsheets; sort by impressions to find big-spender ads
  • JSONL — line-delimited, easy to stream into a warehouse
  • API — fetch programmatically via GET /v2/datasets/{id}/items; webhook on ACTOR.RUN.SUCCEEDED for live pipelines

💰 Pricing

Pay-per-event. You pay for what you get, nothing for what you ask for:

EventPriceWhen charged
actor-start$0.005Base fee, once per run (warm-up, cookie handshake, proxy resolution)
ad-result$0.003Per ad creative written to the dataset

Examples:

PullCost
100 ads$0.31
1 000 ads$3.01
10 000 ads$30.01
100 000 ads (monthly competitor sweep)$300.01

Compare to: building this in-house is roughly two engineer-weeks plus the ongoing cost of maintaining a proxy pool and the TLS-fingerprint replay loop. We've already done it, and we keep it running as Google's RPC shifts underneath us.

🚧 Limitations

  • Region is metadata, not a filter. Google's SearchCreatives RPC ignores the geo target — we confirmed this empirically (see scripts/recon/FINDINGS.md). The Transparency Center browser UI shows a region selector, but the server returns the same creative set regardless. We expose region so you can tag exports by intended locale, nothing more.
  • No region-only browsing. You must supply a searchDomain or advertiserId. There is no "all ads in country X" mode on the public RPC. If Google adds one we'll wire it in.
  • Video / rich creatives return a content.js URL, not an MP4. Rendering the actual video frame requires executing Google's JS bundle — out of scope for v1.
  • Date range is enforced by Google, not us. They retain roughly 12 months of history. Requesting older dates clips to that window.
  • Large advertisers hit pagination caps. Google's infrastructure stops responding past roughly 1 000 ads per query. Nike's library claims ~300 000 ads; the default maxPages=25 is intentionally conservative. Raise it for full-history pulls knowing you may hit the server-side ceiling.

❓ FAQ

Is this legal?

Yes. The Google Ads Transparency Center is a public registry Google operates under EU DSA and US regulatory pressure. We scrape only what the public UI exposes at a polite cadence, and we do not bypass authentication. We also do not collect personal data — only advertiser-level metadata.

Does Google have an official API for Ads Transparency data?

No. As of 2026, Google publishes no official API for the Transparency Center. We reverse-engineered the internal SearchCreatives RPC, replay it with a real-browser fingerprint, and keep the implementation current as the endpoint evolves. The "google ads transparency api" you may have searched for is exactly what this Actor provides.

Is this the same thing as a "Google Ads Library" scraper?

Effectively yes. Google doesn't officially brand it "Ads Library" the way Meta does — the product is the Google Ads Transparency Center — but that's the phrase people search for when they mean the same target. Same registry, same data, this Actor covers it either way.

Why replay an RPC instead of automating a browser against the UI?

Speed and stability. A browser-automation approach re-renders the Transparency Center page for every batch of creatives and re-parses whatever DOM structure Google shipped that week. Talking to the internal RPC directly returns ~40 structured creatives per round-trip and doesn't break when Google reshuffles frontend markup — see Why we replay the RPC above.

Why is the region selector marked "display only"?

Because we empirically confirmed the RPC ignores it. Other scrapers on the Store claim region filtering; we tested every plausible RPC body shape and none returned a region-narrowed result set. We would rather under-promise than ship broken filtering. If Google adds a server-side region filter, we will wire it in immediately.

Why isn't there a search-by-keyword mode?

Google's RPC does not expose one. You search by advertiser. For brand-keyword monitoring, give us the domain (e.g. nike.com) and the scraper returns every ad pointing at that domain — including those bought by competitors bidding on your name.

Can I scrape political ads specifically?

Not yet — political ads live in a separate Google library with its own endpoints. Open an issue on the Apify Store listing if you want this; we will prioritize based on demand.

How do I export to Google Sheets or a database?

Three options:

  1. Console → Storage → Export for one-off CSV downloads.
  2. Webhook the dataset URL to a Make / Zapier flow that appends to Sheets.
  3. Apify integration nodes in Airbyte, n8n, or your warehouse loader.

Some preview URLs are null. Why?

Rich, video, and animated creatives expose only a content.js URL — Google renders the preview via JavaScript. Static image creatives give you a direct preview_image_url. If you need actual video frames, post-process the content.js URL with a headless browser downstream.

The number of returned ads is less than Google's reported total. Why?

Google paginates and stops responding past an internal limit we have observed at roughly 1 000 ads per query. For very large advertisers, raise maxPages beyond the default of 25 if you need fuller coverage.

How do I scrape google ads transparency center data automatically on a schedule?

Go to Apify Console → Schedules, attach this Actor, and set your cron. Weekly is the right cadence — Google updates the Transparency Center daily at most. More frequent polling wastes credit without new signal.

What integrations does this Actor support?

  • Schedule — Apify Console → Schedules tab → run weekly for monitoring.
  • Webhooks — register ACTOR.RUN.SUCCEEDED to fire your downstream pipeline as soon as the dataset is final.
  • APIPOST /v2/acts/DevilScrapes~google-ads-transparency/run-sync-get-dataset-items returns the full result set in one synchronous call (good for up to a few thousand ads).
  • Make / Zapier — every Apify Actor surfaces as a node out of the box.
  • n8n — use the Apify community node; a workflow template is available on n8n.io/workflows.

💬 Your feedback

Spotted a bug, missing field, or want a new feature? Open an issue on the Apify Store listing — we read every one.

Built by Devil Scrapes — Apify Actors with attitude. PPE, transparent pricing, no junk fields.