~ / guides / Is Scraping TikTok Legal? Terms of Service & robots.txt Explained

Is Scraping TikTok Legal? Terms of Service & robots.txt Explained

AQ
Aria Quinn
TikTok data engineer · about the author
the short version
  • Scraping publicly visible TikTok data is generally treated as legal in the US after hiQ v. LinkedIn and Meta v. Bright Data. The federal computer-crime law (the CFAA) does not cover public pages.
  • It still breaks TikTok's Terms of Service. Section 3.4 (updated January 22, 2026) bans scraping by any automated system or bot without written approval. That is a contract risk, plus an access risk: blocks, CAPTCHAs, account bans.
  • TikTok's robots.txt blocks AI crawlers by name (GPTBot, ClaudeBot, CCBot, Bytespider) and disallows paths like /api/ and /search?, while allowing /foryou, /@user and /tag.
  • Private data, logged-in content, and personal data are a different matter. GDPR and CCPA apply to personal data regardless of any US court ruling, and copyright protects the videos themselves.

I spent a morning doing the boring part of this question properly: I read TikTok’s actual Terms of Service, pulled its live robots.txt, and went through the scraping case law (hiQ, Meta v. Bright Data) instead of trusting a forum thread. The short version is that “is scraping TikTok legal” has two answers that people constantly mix up, and keeping them separate is the whole game.

One answer is about the law: can a court punish you for it. The other is about the contract and the technical wall: TikTok’s Terms of Service and its blocking systems. You can be on the right side of the law and still be breaking TikTok’s terms and getting your requests blocked. Below I walk through each layer with the exact wording and the dates, then where TikTok Shop, personal data, and copyright change the picture.

Scraping publicly accessible TikTok data is generally legal under US federal law, because the main computer-crime statute does not reach data that anyone can view without logging in. The data-scraping legality question turns on two court rulings. Neither involved TikTok directly. Both set the rules every platform now operates under.

In hiQ Labs v. LinkedIn, the Ninth Circuit held that the Computer Fraud and Abuse Act (CFAA), the federal “unauthorized access” law, does not apply to scraping data that is publicly available on the open web. The court reaffirmed this in April 2022 after the Supreme Court sent the case back for another look (the Ninth Circuit opinion is on CourtListener). The logic: if a page is open to the public, accessing it with a bot is not “breaking in,” so the anti-hacking statute is the wrong tool.

The second case is more recent and more on point for a social platform. In Meta Platforms v. Bright Data, decided January 23, 2024, Judge Edward Chen of the Northern District of California granted summary judgment to the scraper, with the opinion stating plainly that the “Facebook and Instagram Terms do not bar logged-off scraping of public data.” The reasoning matters: when Bright Data scraped while logged out, it was not a “user” bound by Meta’s terms at all, it was an ordinary visitor. Eric Goldman’s Technology and Marketing Law Blog has the full write-up and the opinion, and Meta dropped the case the following month.

So for TikTok, public profile pages, public video metadata, hashtags, and view counts sit on the protected side of that line under US law. That is the answer to the federal-law question. It is also where most people stop reading, and that is the mistake, because TikTok’s Terms of Service are a separate layer with their own teeth.

Does TikTok’s Terms of Service prohibit scraping?

Yes, TikTok’s Terms of Service explicitly prohibit scraping by automated means, and this is a contract you agree to, separate from any law. I pulled the current US terms, last updated January 22, 2026. Section 3.4, titled “What you can’t do on the Platform,” contains the operative sentence. You may not:

“scrape, crawl, export or otherwise extract any data or content in any form, for any purpose, from the Platform using any automated system or software, including automated ‘bots,’ except as approved in writing by TikTok USDS Joint Venture”

That language is broad on purpose. “Any data or content in any form, for any purpose” covers public and private alike, and “any automated system or software” covers your Python script, a headless browser, and a no-code tool equally, with the only carve-out being prior written approval. The same section also bans reverse engineering the platform, “including its algorithms, code, or infrastructure.” You can read the full Terms of Service on TikTok’s legal pages; TikTok’s separate Developer Terms of Service carry their own scraping prohibition for anyone building on the official APIs, so the ban on automated extraction without written permission applies on both the consumer and developer sides.

Here is the distinction that trips people up, laid out directly:

LayerWhat it governsWho enforces itTypical consequence
US federal law (CFAA)Whether public-data scraping is a crimeFederal courtsGenerally does not apply to public data
TikTok Terms of ServiceWhether you broke your contract with TikTokTikTok, via civil claimAccount ban, cease-and-desist, breach-of-contract suit
Privacy law (GDPR, CCPA)Handling of personal dataRegulators, EU and CaliforniaFines, regardless of CFAA status
CopyrightReusing the videos and content themselvesRights holdersTakedown, infringement claim

The Terms-of-Service row is what changes a “legal” scrape into a risky one. Meta v. Bright Data helps here, because that ruling suggests a logged-off scraper may not be bound by the terms in the first place, but that question is decided court by court and TikTok’s wording was written to argue the opposite. Treating the terms as a live constraint is the safe assumption. The clearest way to stay inside them is to use access TikTok actually sanctions, which is its official APIs, the subject of a later section.

The legality of scraping TikTok Shop data follows the same public-versus-private split as the rest of the platform, with one extra wrinkle in the robots file. Product listings, prices, ratings, and seller pages that load for any logged-out visitor are public data, so the US case law that protects public scraping covers them the same way it covers a public profile.

The wrinkle is that TikTok’s robots.txt singles out the Shop. The catch-all User-agent: * block disallows /shop/view/product/ specifically, alongside /search? and /api/. So while the data is public in the legal sense, TikTok has formally asked crawlers to stay off the product detail path. Scraping it stays legal under US public-data rulings. The disallow line still strengthens TikTok’s “you ignored our stated wishes” argument in any Terms-of-Service dispute, and it signals that the Shop paths are watched.

For commercial use of TikTok Shop data, there is a third factor on top of law and contract: the seller and product information can include personal data of individual sellers, which pulls in the privacy rules I cover below. The robots file is the bridge to the next question, because reading it line by line is the clearest signal of what TikTok actively wants left alone.

What does TikTok’s robots.txt say about scraping?

TikTok’s robots.txt does not ban all scraping. It blocks AI crawlers by name and disallows specific paths for everyone else, while leaving the main public surfaces open to general crawlers. I pulled the live file at tiktok.com/robots.txt; here is what it actually contains.

The most aggressive rule targets AI training and assistant bots. These user-agents get a blanket Disallow: /:

CategoryUser-agents blocked by name
AI training crawlersGPTBot, CCBot, Google-Extended, Applebot-Extended, anthropic-ai, ClaudeBot, Bytespider, AI2Bot, meta-externalagent
AI assistant / searchOAI-SearchBot, ChatGPT-User, PerplexityBot, Perplexity-User, Claude-User, Claude-SearchBot, DuckAssistBot, Gemini-Deep-Research, Google-NotebookLM, MistralAI-User
Chinese search enginesBaiduspider, 360Spider, Sogouspider, Yisouspider, PetalBot

For every other crawler, the User-agent: * block draws a path-level line:

Allowed pathsDisallowed paths
/foryou, /discover, /about, /legal, /tag, /music, /share, /transparency, /community-guidelines/api/share/settings, /api/recommend/embed_videos, /search?, /search/live?, /inapp, /auth, /embed/@, /link, /shop/view/product/, /discover/trending/detail/

Two things stand out. The public content surfaces (the For You feed at /foryou, hashtag pages at /tag, profile shares via /share) are explicitly allowed for general bots, while the internal API endpoints (/api/...), search query strings, and the Shop product path are disallowed. A robots.txt is a request from the site, with no technical lock and no force of law behind it, but ignoring a path the site explicitly disallowed is exactly the conduct that looks bad in a Terms-of-Service fight. Robots rules tell you TikTok’s stated wishes; the blocking systems are what enforce access in practice, and those are stricter than the file suggests.

What happens technically when you scrape TikTok?

Technically, TikTok actively fights scrapers with rate limiting, CAPTCHAs, and behavioral bot detection, so the practical risk is getting blocked long before any legal question arises. TikTok published a post titled “How We Combat Unauthorized Data Scraping of TikTok” (October 25, 2024) describing its approach in its own words.

According to that TikTok privacy post, its named defenses include:

This is where most of the real TikTok scraping legal issues start for an engineer: an automated scraping run hits these defenses on day one, long before anyone reads a contract. TikTok’s scraping policy, stated across the Terms of Service and this post, is that automated extraction without written permission is unauthorized, and the API terms on the developer side say the same for anyone building on the platform.

TikTok also states directly that “unauthorized data scraping violates TikTok’s Terms of Service,” tying the technical wall back to the contract layer. In practice this means a naive scraper hitting public TikTok pages from a datacenter IP runs into challenge pages and HTTP errors quickly, well before TikTok ever considers a legal response. The realistic order of consequences runs technical first, contractual second, legal last:

LikelihoodConsequenceTrigger
Very commonIP block, CAPTCHA wall, HTTP 4xxVolume, datacenter IPs, missing headers
CommonAccount suspensionScraping while logged in, fake accounts
UncommonCease-and-desist letterLarge-scale or commercial scraping noticed by TikTok
RareBreach-of-contract lawsuitSustained, high-profile, commercial use

That rare bottom row is not zero. In the hiQ matter, the scraper ultimately faced a 500,000 US dollar judgment for breaching LinkedIn’s user agreement, as covered in the law-firm analyses of the December 2022 consent judgment. Most projects never get near a courtroom, but they do hit the blocking wall on day one, which is why the sanctioned routes matter.

Personal data and copyright sit outside the public-versus-private scraping debate and follow their own rules, which apply even when the data is public and even when US courts say scraping is fine. This is the layer that catches commercial projects off guard.

On personal data, the EU’s General Data Protection Regulation (GDPR) and California’s Consumer Privacy Act (CCPA) both regulate the collection and processing of information about identifiable individuals. A username, profile photo, bio, and follower list can all be personal data. The fact that a US court would not treat scraping a public profile as a CFAA crime does not exempt you from a regulator’s data-protection rules if you are collecting EU or California residents’ personal data. The two regimes operate on different tracks, and a clean answer on one says nothing about the other.

On copyright, the videos, audio, and images on TikTok are protected creative works, the intellectual property of their creators and TikTok. Collecting public metadata (view counts, captions, hashtags, timestamps) is one thing. Downloading and republishing the videos themselves outside the platform is a copyright question, and “it was publicly viewable” is not a defense to infringement. If your use case touches the media files themselves, beyond the structured data around them, that is the framework to evaluate.

The practical reading across all of this: aggregate, non-personal, public signals (trends, hashtag volume, public view counts) are the lowest-risk data to collect. Targeted collection of individuals’ personal details, or redistribution of the videos, climbs the risk ladder fast. Once you know which side of that line your project is on, the next decision is which access method keeps you there, and TikTok offers a sanctioned one.

How do you collect TikTok data legally and within the terms?

The lowest-risk way to collect TikTok data is through access TikTok sanctions, which means its official APIs for permitted use cases, or a scraper API that pulls only public data for you. Both keep you clear of the “automated access without permission” problem in different ways.

TikTok runs two official programs. The TikTok Research API gives approved academic and nonprofit researchers structured access to public data under its own Research Tools terms, while the Display API and broader developer platform cover app integrations with user consent. These are the routes TikTok explicitly authorizes, so they sidestep Section 3.4 entirely, but their limit is eligibility and scope: the Research API gates on who you are and what you study, the developer APIs are built around logged-in user consent for app integrations, and neither is designed for bulk collection of public data by a general user.

When your use case is collecting public data and you do not qualify for the Research API, a scraper API is the common middle path. It accepts a public URL or identifier and returns parsed data, handling the proxy rotation and challenge-solving on its end so your project stays on public pages and never logs into accounts. In my own testing across the public TikTok endpoints, this is the request shape that maps cleanly to the public profile surface (the same /share and /@user data the robots file allows). Using ChocoData, a request looks like this:

# Public TikTok profile data. Get a key at https://app.chocodata.com/sign-up
curl "https://chocodata.com/api/v1/tiktok/profile?username=nasa&api_key=$CHOCO_API_KEY"

That call targets a public profile, the kind of page TikTok’s own robots.txt permits general crawlers to reach. It does not log in, create fake accounts, or touch the disallowed /api/ and /search? paths, which keeps it on the public-data side that hiQ and Meta v. Bright Data protected. The same approach extends across the public surfaces I cover in my other guides:

None of this makes a scraper API a license to ignore the layers above. It keeps you on public data and out of logged-in territory, which is the part you control. The privacy and copyright questions still depend on what you collect and what you do with it. If you want the hands-on mechanics next, my walkthroughs on how to scrape TikTok and how to scrape TikTok with Python show the actual requests, and the best TikTok scrapers in 2026 compares the tools that handle the blocking for you.

FAQ

Is scraping TikTok legal in the US?

Scraping publicly accessible TikTok data is generally legal under US federal law. In hiQ v. LinkedIn the Ninth Circuit held the Computer Fraud and Abuse Act does not bar collecting public data, and in Meta v. Bright Data a federal judge ruled that logged-off scraping of public data does not breach the platform's terms. Scraping private data, logged-in content, or personal data is a separate question governed by privacy law and TikTok's contract.

Does TikTok's Terms of Service prohibit scraping?

Yes. TikTok's Terms of Service, Section 3.4, prohibit you from using any automated system or software, including 'bots', to scrape, crawl, export, or otherwise extract data from the platform without written approval from the TikTok USDS Joint Venture. Breaking the terms is a contract issue between you and TikTok, separate from whether scraping breaks any law.

What does TikTok's robots.txt disallow?

TikTok's robots.txt fully blocks AI training and assistant crawlers by name (GPTBot, ClaudeBot, CCBot, Google-Extended, Bytespider, PerplexityBot and others). For everyone else it disallows paths such as /api/, /search?, /inapp, and /shop/view/product/, while allowing /foryou, /@username via /share, /tag, and /discover.

Is scraping TikTok Shop data legal?

The legality of scraping TikTok Shop data follows the same public-versus-private line. Product listings, prices, and seller pages that anyone can view without logging in sit on the public side that US courts have protected. The catch is that TikTok's robots.txt specifically disallows /shop/view/product/, and the Terms of Service ban automated extraction, so Shop scraping carries the same access and contract risk as the rest of the platform.

Can scraping TikTok get me sued or banned?

A lawsuit over public-data scraping is uncommon, but the contract claim is real: in the hiQ case the scraper ended up with a 500,000 US dollar judgment for breach of LinkedIn's user agreement. The far more frequent outcome is technical: rate limiting, CAPTCHAs, IP blocks, and account suspension. TikTok states it uses all of these to combat unauthorized scraping.

AQ
Aria Quinn
I've built TikTok data pipelines for years. On tiktokscraperapi.com I run TikTok scraping methods against live pages and publish what actually holds up.