Back to Blog
Engineering

Instagram Scraper: Open-Source Python Alternative to Apify

Scrape Instagram profiles, posts, reels and comments with open-source Python code. Run it on bro’s cloud browsers as an alternative to Apify.

Instagram Scraper: Open-Source Python Alternative to Apify

TL;DR: After proving bro can bypass the toughest anti-bot systems, we built an open-source Instagram scraper.

It matches Apify's capabilities - extracting profiles, posts, reels and comments for a fraction of the cost. Check out the source code below.

Motivation

Our previous benchmark proved that bro can passively bypass the most aggressive bot protections in the web.

But getting in is only half the battle. For our next experiment, we wanted to build something you can actually use.

We set our sights on Apify's Instagram Scraper. It's a gold standard in the industry and we wanted to see if we could build an open-source alternative that matches its data extraction quality but runs cheaper and gives developers full control over the code.

Because it runs on bro, you never have to worry about browser infrastructure, proxy rotation or anti-bot bypassing. You just run the code and get the data.

What data can it extract?

Our Python-based scraper can run directly from your CLI or be integrated into your own backend.

0:00 / 0:00

Feed it a list of URLs, usernames or search queries so it'll pull:

TargetWhat you get
ProfilesBio, follower & following counts, profile pictures and account details
Posts & ReelsCaptions, dates, engagement stats, media URLs and carousel items
Comments & RepliesText, authors, timestamps, likes and full reply threads
StoriesActive story items, timestamps and media URLs
Tagged PostsPosts where the target account is tagged
Search & LocationsHashtags, location details and associated posts

You can export the results as JSON, JSONL or CSV.

To make migrating easy, we kept familiar settings like directUrls, resultsType and resultsLimit. If you already use an Apify-style input payload, it'll work great out of the box.

How it actually works

Instead of forcing you to setup dozens of Chrome docker containers, fight anti-bot systems or manage proxies, bro runs the browser sessions remotely. Your Python script just tells it what to do.

Pipeline
Set inputs and limits
bro navigates to the target IG pages
Scraper parses logs from internal APIs
auto-saves progress
exports JSON/CSV

Here is the secret sauce: instead of visually rendering the page and scraping HTML elements one by one, our scraper looks for the structured data Instagram already uses under the hood. By intercepting these internal API responses, we grab the exact data we need making the process faster, more reliable and significantly cheaper.

If that optimal route isn't available, the scraper seamlessly falls back to reading the rendered page content or using optional AI data extraction.

No more lost data

Anyone who scrapes at scale knows the pain: your scraper runs perfectly for a couple of hours, hits an unexpected error, crashes and you lose everything.

We fixed that. Our scraper features a resumable collection mode. It saves records and pagination state locally as it goes.

If your job gets interrupted, you just restart it and point it to the saved run directory. The scraper checks its checkpoint, ignores duplicates and picks up exactly where it left off. You keep the data you already paid to extract.

Why bro makes this cheaper?

Building scraper on bro made it possible to reduce costs up to 75x

Building scraper on bro made it possible to reduce costs up to 75x

Building this on bro gave us total control over the browser session. We can execute JavaScript, inspect network requests and responses with just a few lines of code.

In a single benchmark run, we collected ~16k comments (deep comments + replies) and the total compute cost for that session was just $0.4824. Which gives us $0.03 per 1k comments (75x cheaper than Apify).

We'll release a dedicated blogpost with side by side comparison soon!

The biggest challenge? Comments extraction.

At first, we could only collect 15 comments before the page stopped loading more. Eventually, we managed to completely bypass the UI limitations by inspecting the page's network logs to collect ~16k comments from 12 posts.

Our scraper ends up being incredibly cost-effective compared to the similar solutions due to:

  1. Billing - bro charges only for the raw resources your session consumes (compute time and proxy bandwidth) rather than per post / profile / reel / etc.
  2. Session configuration - bro lets you explicitly set your proxy tier and traffic routing policy. You have full control over your bandwidth costs and success rates.

Source code

We've open-sourced the entire project: the source code is on GitHub

All you need is Python 3.10+, a bro API key and your Instagram cookies if you plan to scrape session-dependent data (like deep comment threads or stories).

You still can scrape profile, reels and post data without utilizing cookies.

The project has been built for educational purposes only.

Integrate to your project with just one prompt

Focus on the data, not the infrastructure.
Copy the prompt and paste it to your AI agent.

Get API key
No commitment
No credit card required
Pay as you go

FAQ

What can this Instagram scraper collect?

Do I need to provide Instagram cookies?

Do I have to use bro to run this scraper?

Do I need to run a browser locally?

Does it replace Apify completely?

What does it cost to run?