Content
<div align="center">
# Tool List
> "Have you studied 10 benchmarking bloggers? Can you clearly explain the winning strategy of any one of them?"
You've flipped through 50 notes, but still can't explain why they went viral?<br>
You've used AI to write imitation posts, but they only feel generic - it doesn't know who you're trying to learn from.<br>
You've switched to another benchmarking blogger, and everything starts from scratch again?<br>
You've also posted thirty notes yourself, but every time you create, you start from scratch?<br>
**Distill the winning strategy of any blogger (including yourself) from real notes, and put it into your AI.**
Enter the blogger's account, and get their winning strategy in 30 minutes -
put it into your AI, and it becomes your permanent content coach.
[](https://www.python.org/downloads/)
[](LICENSE)
[Output](#output) · [Analysis Mode](#analysis-mode) · [Get Started in 30 Seconds](#30-second-get-started) · [API Configuration](#api-configurationtikhub) · [Cost Estimation](#cost-estimation) · [Join the Discussion Group](#join-the-discussion-group) · [Find the Author](#find-the-author)
</div>
---
<details>
<summary><b>⚠️ Read Before Use: Disclaimer and Data Processing Instructions</b> (click to expand)</summary>
This tool is for **learning and research** purposes only. Before using, please confirm that you understand the following points:
- **Data Source**: Public data is obtained through [TikHub](https://tikhub.io)'s public REST API (no login simulation, no cookie injection, no encrypted interface cracking).
- **Scope of Application**: Only for research on **publicly published** content (public accounts, public notes/posts, public comments). Please do not use it for crawling private content, bypassing platform risk control, or any behavior that violates the User Agreement of Xiaohongshu, Douyin, and relevant laws and regulations.
- **Commenter Privacy**: This tool will automatically anonymize commenters' identities as "Reader1 / Reader2 / Author", without saving commenter nicknames, user IDs, avatars, or IP locations. Comment content will be retained for content research.
- **Commercial Use**: If you need to use it for commercial purposes, please confirm whether Xiaohongshu/Douyin platform policies and TikHub service terms allow it. The author does not endorse the compliance of commercial use.
- **Risk Assumption**: You are solely responsible for all consequences (including but not limited to account risks, data accuracy, and third-party API fees) of using this tool.
**Compliance Boundaries** (refer to "Data Security Law", "Personal Information Protection Law", "Anti-Unfair Competition Law"):
✅ These are okay: Personal learning / demand analysis / market research; only crawling publicly accessible content without login; using data only for your own research, not for external dissemination or monetization.
❌ These are risky: Selling or distributing others' data; crawling content that requires login; using large-scale crawled data as the core data source for commercial products.
See [DISCLAIMER.md](./DISCLAIMER.md) for detailed terms · Security strategy see [SECURITY.md](./SECURITY.md)
</details>
---
## What It Does
Turn "manual note flipping → feeling-based summary" into "one sentence → automatic distillation → get results".
```
Deconstruct/Distill Blogger XX (best with a link🔗)
```
AI completes the entire process automatically:
```
Environment check → endpoint automatic detection → note collection → statistical analysis → cognitive layer extraction → distillation → report generation
```
You get not just raw data, but two things (or three with transcription enabled):
| Output | For Whom | What to Do |
|--------|------|--------|
| **Creation Guide Skill Folder** | For Your AI | Put into AI tools, use XX blogger's cognition, strategy, and content style to help you create or diagnose accounts |
| **HTML Distillation Report** | For You | Open in a browser, quickly read the blogger's persona, strategy, and winning patterns |
| **Transcription Text** (optional) | For Your AI | Turn video speech into text, supplement analysis data, and let AI read the blogger's real speech content |
---
## Output
### Creation Guide Skill (for AI)
Generate an installable Skill folder. Put it into your AI tool:
```
.claude/skills/XX Blogger_Creation_Guide.skill/
└── SKILL.md
```
After installation, in a new conversation, directly say:
```
Write a Xiaohongshu note about AI tool recommendations in the style of XX blogger
```
AI writes according to the distilled formula, tone, and structure - not generic "lively style", but specific rules extracted from real notes.
**Three-layer distillation structure:**
| Layer | Answer What | Example Content for Skill |
|------|---------|----------------------|
| **Cognitive Layer** | What does TA think? | "TA believes 'anti-consensus has communication power', 80% of viral content uses unconventional perspectives" |
| **Strategy Layer** | How does TA operate? | "Posting every Wednesday at 8 PM, a hidden like ratio of 0.85 indicates fans are saving as tutorials" |
| **Content Layer** | How does TA write? | "35% digital titles, 25% interrogative titles, always starting with personal stories" |
8 chapters: instructions → cognitive layer → strategy layer → content layer → creation forbidden zones → comparison examples → topic inspiration → limitation self-inspection
### HTML Distillation Report (for You)
Single-file HTML, open in a browser to read.
| # | Module | Content |
|---|------|------|
| 1 | At-a-Glance | Followers/likes/viral rate/bottom formula, one card summary |
| 2 | Persona Deconstruction | Who is TA, what does TA represent, why do fans follow TA |
| 3 | Cognitive Layer | Core beliefs × viewpoint tension × value stance |
| 4 | Strategy Layer | Series planning × hot topic hijacking × operational rhythm |
| 5 | Top 10 Viral Content | Dissect each: why it went viral? what can be learned? |
| 6 | Content Formula | 11 title formulas × opening templates × CTA strategies |
| 7 | Topic Inspiration | Top 15 topic directions, sorted by difficulty × potential |
| 8 | Data Panel | Avg likes/avg saves/save-like ratio/video vs graphic/posting frequency |
| 9 | Development Trend | Early vs recent changes, transformation path (with confidence level) |
| 10 | Core Conclusion | Three-sentence summary + actionable suggestions |
### Transcription Text (optional)
After distillation, say "extract transcription", Whisper automatically transcribes video audio into text:
- File saved as `{Blogger Name}-Transcription.md`, can be directly given to AI for in-depth analysis
- Default Base model, accuracy about 70%; can upgrade models in Skill, each upgrade doubles transcription time
- Only supports videos within 10 minutes, longer videos skipped
---
## Analysis Mode
| Mode | What You Want to Do | Output |
|------|-----------|--------|
| **A — Deconstruct Benchmark Blogger** | Learn TA's strategy | HTML Report + `{Blogger Name}_Creation_Guide.skill/` |
| **B — Distill Your Own Account** | Build your own creation workflow | HTML Report + `{Your Username}_Creation_Gene.skill/` |
Note quantity options: 30 notes (quick) / 50 notes (recommended) / 80 notes (in-depth)
---
## Extended Play
After the first distillation, say a trigger word to unlock two advanced analyses - no need to re-collect data, directly run based on existing data:
| Play | Trigger Word | Platform | Description |
|------|--------|------|------|
| 🎨 Cover Visual Style Analysis | "Analyze Cover" | Both platforms | 8-dimension visual dissection: cover style type / composition / title hook / text design / character appearance / decorative elements / information density / visual consistency, output cover formula + content matching degree assessment + 3 suggestions (no extra API calls) |
| 📈 Keyword Trend Insight | "Keyword Trend" | Both platforms | Analyze blogger's core keyword heat trend and audience portrait, mining content direction opportunities |
---
## 30-Second Get Started
```bash
git clone https://github.com/otter1101/blogger-distiller.git
cd blogger-distiller
python install.py
```
In your AI programming assistant, say:
```
Deconstruct Xiaohongshu/Douyin blogger XX
```
The system will ask you to select mode (A or B), platform (Xiaohongshu / Douyin), and note quantity, then fully automatic.
### Skill Installation (Advanced)
**Method 1: Let AI Help You Install (Recommended)**
Open your AI Agent (local/cloud), say:
```
Please help me download https://github.com/otter1101/blogger-distiller's skills and tell me how to use them
```
AI will automatically complete the download and configuration.
**Method 2: Manual Installation**
Download the Skills zip package, extract to the corresponding folder, and complete the installation according to the prompts.
---
## API Configuration (TikHub)
This tool uses [TikHub](https://tikhub.io)'s REST API to fetch Xiaohongshu/Douyin data. **Entirely through external API calls, no simulated login, no cookie injection, zero account risk.** This is currently the lowest-cost, most comprehensive third-party API solution. The author has no interest in TikHub and does not profit from it, only providing a usable solution.
You need to register a TikHub account, recharge, and enable API permissions:
### Step 1: Register
Visit [https://user.tikhub.io](https://user.tikhub.io) to register and complete email verification.
### Step 2: Recharge
Log in and recharge in the **Console → Recharge/Package** page. Recommended pay-as-you-go, charged by actual call count.
### Step 3: Enable Permissions
In **Console → API Permissions**, **select all Xiaohongshu (xiaohongshu) and Douyin (douyin) related endpoints**. The system will automatically detect which endpoints are available during startup and does not require manual selection. The more endpoints enabled, the stronger the automatic fault tolerance.
### Step 4: Generate API Token
In the **Console → API Token** page, generate a Token key, then give it to your agent!
ps: If the agent cannot parse the API key, try 3 solutions:
1. Retry/check network issues
2. Change the key, no need to recharge
3. Change an agent, claudcode/codex/workbuddy/openclaw/hermes all work
<details>
<summary>Optional: Custom RPS Acceleration</summary>
TikHub has different RPS (requests per second) limits for different packages. Default RPS=10. If your package has a higher RPS, you can manually accelerate:
```bash
export TIKHUB_RPS=20 # halve the interval, double the speed
```
</details>
---
## Cost Estimation
Approximate cost of distilling one blogger (based on TikHub pay-as-you-go):
| Note Quantity | Estimated Cost |
|----------|---------|
| **30 Notes** (quick) | ¥1 ~ 3 |
| **50 Notes** (recommended) | ¥2 ~ 5 |
| **80 Notes** (in-depth) | ¥4 ~ 8 |
> Actual cost depends on comment quantity, endpoint availability, and retry times. Based on real usage statistics, please refer to [TikHub Official Pricing](https://tikhub.io/pricing) for accuracy. The system has built-in multiple saving strategies (startup detection to avoid invalid endpoint requests, session cache for dead links, etc.) to minimize invalid API calls.
---
## Architecture Design
**'Script as the lower bound, AI as the upper bound'**
| Role | Proportion | Responsible for |
|------|------|---------|
| **Script** | 30% | Data collection, statistics calculation, title pattern recognition, CTA extraction, save-like ratio calculation |
| **AI** | 70% | Belief extraction, cognitive tension discovery, persona deconstruction, causal analysis, personalized suggestions |
**Core Features:**
- **Zero Configuration** — Automatic dependency download, service startup, ready to use
- **Endpoint Automatic Detection** — Automatically detect API endpoint availability during startup and dynamic sorting, adaptive + adjustment
- **Multi-Endpoint Fallback** — 4 endpoint pools (web_v2 / app / web_v3 / app_v2) automatically degrade, one down, three still available
- **Cross-Domain Universal** — No content type preset, dynamic clustering based on actual note content, applicable to any niche
- **Breakpoint Recovery** — Automatic save every 10 notes, no data loss in case of network interruption or interruption
---
## Project Structure
```
blogger-distiller/
├── SKILL.md # Skill definition (read by AI programming assistant)
├── run.py # One-click operation entry
├── install.py # Automatic installation script
├── scripts/
│ ├── crawl_blogger.py # Phase 1: Collection entry routing (automatically distribute to each platform)
│ ├── crawl_xhs.py # Phase 1-XHS: Xiaohongshu data collection
│ ├── crawl_douyin.py # Phase 1-DY: Douyin data collection
│ ├── crawl_common.py # Shared tools for collection layer (directory/JSON/progress/rate limiting)
│ ├── analyze.py # Phase 2: Data analysis (clustering + tagging + TOP10 + cognitive layer)
│ ├── deep_analyze.py # Phase 3: Data draft + AI distillation task generation
│ ├── verify.py # Phase 4: Data verification (both platforms)
│ └── utils/
│ ├── tikhub_client.py # TikHub API client (multi-platform routing + rate limiting + degradation)
│ ├── endpoint_router.py # Endpoint pool routing + automatic degradation engine
│ ├── xhs_endpoints.json # Xiaohongshu endpoint pool (4 groups × 7 categories = 28 endpoints)
│ ├── douyin_endpoints.json # Douyin endpoint pool (10 functional pools = 19 routes)
│ ├── adapters.py # Response data normalization adapter (XHS + Douyin)
│ ├── cover_analyzer.py # Cover visual style analysis (extended play: card A)
│ ├── index_client.py # Keyword trend insight (extended play: card B, both platforms)
│ ├── common.py # Platform registry + shared tool functions
│ ├── privacy.py # Data anonymization (both platforms)
│ ├── first_run.py # First run compliance confirmation
│ ├── quality.py # Data quality inspection tool
│ └── transcript.py # Transcription extraction (Whisper integration, Xiaohongshu + Douyin)
```
---
## Version History
| Version | Milestone |
|------|--------|
| v0.1–v1.0 | SKILL design → Full phase script → Universal refactoring → Full automatic environment preparation → Input end stabilization (scan login version, deprecated) |
| v1.7 | Cognitive layer extraction (view sentence / thinking mode / value word) |
| v1.8 | Output refactoring (HTML report + Skill folder) |
| v1.9 | Full process switching to new paradigm + collection bug fix |
| **v2.0** | **API + compliance transformation release** |
| **v2.2** | **Multi-platform extension**: added Douyin full-chain collection (22 endpoints), dual-platform analysis engine, cover visual analysis (card A), keyword trend insight (card B), atomic card extension |
| **v2.3** | **Transcription extraction**: integrated Whisper model, Xiaohongshu and Douyin video can be transcribed into text with one click (Base model ~70% accuracy / multi-model optional); 10-minute video automatic transcription, longer videos skipped; transcription export as MD |
---
## License & Disclaimer
This project is open-sourced under [MIT License](./LICENSE).
This tool is for **learning and research** purposes only. Read:
- [DISCLAIMER.md](./DISCLAIMER.md) — Disclaimer and usage boundary
- [SECURITY.md](./SECURITY.md) — Data security and privacy handling
If you find this tool useful, please star it in the right corner!
---
## Join the Discussion Group
<div align="center">
<img src="assets/qr-group.jpg" width="260" alt="Blogger Distiller Discussion Group">
</div>
vx:catsanddogs666 (reply not timely during the day)
---
## Find the Author
<div align="center">
<img src="assets/qr-xiaohongshu.jpg" width="220" alt="Xiaohongshu: Aha 水濑"> <img src="assets/qr-douyin.jpg" width="220" alt="Douyin: @Aha 水濑">
</div>
Connection Info
You Might Also Like
markitdown
Python tool for converting files and office documents to Markdown.
OpenAI Whisper
OpenAI Whisper MCP Server - 基于本地 Whisper CLI 的离线语音识别与翻译,无需 API Key,支持...
oh-my-opencode
Background agents · Curated agents like oracle, librarians, frontend...
bm.md
A better Markdown typesetting assistant | One-click adaptation for WeChat...
polymarket-mcp-server
AI-powered trading platform for Polymarket with 45 tools and real-time monitoring.
evo-ai
Evo AI is an open-source platform for creating and managing AI agents.