Key facts
- Metadata for 5.6 billion public TikTok videos was posted on Hugging Face.
- The dataset spans July 2014 through October 2026.
- Data includes captions, hashtags, sound IDs, and engagement counts.
- The data was reportedly obtained by accessing TikTok's private mobile API.
- Commercial use and daily updates are routed to datasocial.ai.
- The scraper's source code is sold separately for $1,699.
A developer known as hashfunction has made a substantial dataset of TikTok video metadata publicly available on Hugging Face, a platform for open-source AI resources. The dataset comprises metadata for approximately 5.6 billion public TikTok videos, covering a period from July 2014 through October 2026. It is offered free of charge under a non-commercial license (CC BY-NC 4.0).
The data includes details such as video captions, hashtags, sound IDs, on-screen text, TikTok Shop product and seller IDs, and engagement metrics like views, likes, comments, shares, saves, and downloads. The entire collection is stored in 460 GB of Parquet files, organized monthly.
According to the developer's write-up, the data was obtained by accessing TikTok's private mobile API, bypassing the need for user logins. This method reportedly involved reverse-engineering request signatures and spoofing device identities. TikTok's terms of service explicitly prohibit automated data scraping without written approval.
While the dataset is free for research and personal use, commercial applications, access to creator profiles, and daily updates are directed to datasocial.ai. The source code used for scraping is also available for purchase separately for $1,699.
This release comes amid ongoing legal battles over data scraping, such as Reddit's lawsuit against Perplexity and other firms for allegedly harvesting content for AI training. The availability of such large datasets is crucial for AI developers seeking to build models that can predict viral content, understand user engagement, and analyze language trends on short-form video platforms.

