前两篇解决了"内容结构化"和"Schema标记",但知识单元如果只躺在自己的网站上,AI大概率不会来抓取。GEO的第三块拼图是信源分发——把你的知识单元主动推送到AI高频抓取的平台上,建立多维度的引用网络。本文提供一套Python工具,自动生成多平台分发内容,并跟踪引用效果。
维度 | 自有网站 | 第三方信源 |
|---|---|---|
AI抓取频率 | 低(新站尤甚) | 高(平台权重高) |
索引速度 | 数天~数周 | 数小时~数天 |
引用可信度 | 需积累 | 平台背书 |
覆盖场景 | 单一 | 多平台交叉验证 |
核心逻辑:AI不是只搜一个地方。它在回答用户问题时,会同时检索多个高权重信源进行交叉验证。你的知识单元出现在越多可信平台上,被AI引用的概率就越大。
平台类型 | 代表平台 | AI抓取强度 | 适合内容 |
|---|---|---|---|
知识库 | Wikipedia、百度百科、维基百科 | ★★★★★ | 定义、事实 |
问答社区 | Quora、知乎、Reddit | ★★★★☆ | 解释、对比 |
技术文档 | GitHub、官方Docs、MDN | ★★★★★ | API、参数 |
新闻媒体 | 主流新闻站、PR发布平台 | ★★★★☆ | 动态、公告 |
专业社区 | Stack Overflow、CSDN、掘金 | ★★★★☆ | 教程、代码 |
社交媒体 | Twitter/X、LinkedIn | ★★★☆☆ | 品牌提及 |
视频平台 | YouTube、B站(字幕可索引) | ★★★☆☆ | 教程、评测 |
学术平台 | arXiv、Google Scholar | ★★★★★ | 研究、论文 |
商业目录 | Crunchbase、企查查 | ★★★★☆ | 公司信息 |
博客平台 | Medium、Substack、微信公众号 | ★★★☆☆ | 深度文章 |
┌──────────────┐ ┌──────────────┐ ┌──────────────┐
│ Knowledge │────│ Distributor │────│ Platform │
│ Units Input │ │ (分发引擎) │ │ Adapters │
│ (知识单元) │ │ │ │ (平台适配器) │
└──────────────┘ └──────┬───────┘ └──────────────┘
│
┌──────▼───────┐
│ Tracker │
│ (效果跟踪) │
└──────────────┘"""
geo_distributor.py - GEO信源分发引擎
技术栈: Python / json / hashlib / typing / datetime / dataclasses
场景: 将原子化知识单元分发到多个AI高频抓取的信源平台
"""
import json
import hashlib
import re
from dataclasses import dataclass, field, asdict
from typing import List, Dict, Optional, Any
from datetime import datetime, timedelta
from enum import Enum
# ==================== 模块1: 数据模型 ====================
class PlatformType(Enum): 5057.baike.tongsou.com
"""平台类型枚举"""
WIKI = "wiki"
QA = "qa"
DOCS = "docs"
NEWS = "news"
COMMUNITY = "community"
SOCIAL = "social"
VIDEO = "video"
ACADEMIC = "academic"
BUSINESS = "business"
BLOG = "blog"
@dataclass
class PlatformConfig: 6004.baike.tongsou.com
"""
平台配置
定义每个信源平台的特性
"""
name: str # 平台名称
platform_type: PlatformType # 平台类型
base_url: str # 平台基础URL
content_limit: int = 5000 # 内容字数限制
supports_structured: bool = False # 是否支持结构化数据
ai_crawl_frequency: str = "daily" # AI抓取频率: hourly/daily/weekly/monthly
authority_score: float = 0.8 # 权威度评分 (0~1)
api_available: bool = False # 是否有API可自动发布
notes: str = "" # 备注
def to_dict(self) -> Dict[str, Any]: 6007.baike.tongsou.com
return asdict(self)
@dataclass
class DistributionItem: 6008.baike.tongsou.com
"""
分发项
一个知识单元在某个平台上的分发记录
"""
unit_id: str = "" # 知识单元ID
entity: str = "" # 实体名称
attribute: str = "" # 属性
value: str = "" # 值
platform: str = "" # 目标平台
platform_type: str = "" # 平台类型
content: str = "" # 适配后的内容
target_url: str = "" # 目标URL(发布后填写)
status: str = "pending" # 状态: pending/published/failed
published_at: str = "" # 发布时间
crawl_detected: bool = False # 是否检测到AI已抓取
reference_count: int = 0 # 被AI引用次数(估算)
def to_dict(self) -> Dict[str, Any]: 14009.baike.tongsou.com
return asdict(self)
@dataclass
class DistributionPlan: 14013.baike.tongsou.com
"""
分发计划
一批知识单元的分发方案
"""
plan_id: str = ""
source_url: str = ""
total_units: int = 0
items: List[DistributionItem] = field(default_factory=list)
created_at: str = ""
summary: Dict[str, Any] = field(default_factory=dict)
def to_json(self, indent: int = 2) -> str: 14068.baike.tongsou.com
return json.dumps({
"plan_id": self.plan_id,
"source_url": self.source_url,
"total_units": self.total_units,
"item_count": len(self.items),
"items": [i.to_dict() for i in self.items],
"summary": self.summary,
"created_at": self.created_at,
}, ensure_ascii=False, indent=indent)
# ==================== 模块2: 平台配置库 ====================
class PlatformRegistry: 14070.baike.tongsou.com
"""
平台注册表
管理所有支持的信源平台配置
"""
# 预定义平台配置
PLATFORMS = [
PlatformConfig(
name="Wikipedia",
platform_type=PlatformType.WIKI,
base_url="https://en.wikipedia.org/wiki/",
content_limit=10000,
supports_structured=True,
ai_crawl_frequency="hourly",
authority_score=0.98,
notes="最高权重信源,适合实体定义类内容",
),
PlatformConfig(
name="百度百科",
platform_type=PlatformType.WIKI,
base_url="https://baike.baidu.com/item/",
content_limit=8000,
supports_structured=False,
ai_crawl_frequency="hourly",
authority_score=0.95,
notes="中文AI模型重要信源",
),
PlatformConfig(
name="知乎",
platform_type=PlatformType.QA,
base_url="https://www.zhihu.com/question/",
content_limit=5000,
supports_structured=False,
ai_crawl_frequency="daily",
authority_score=0.82,
notes="适合问答形式的知识解释",
),
PlatformConfig(
name="GitHub",
platform_type=PlatformType.DOCS,
base_url="https://github.com/",
content_limit=50000,
supports_structured=True,
ai_crawl_frequency="hourly",
authority_score=0.92,
notes="技术类内容最高权重信源",
),
PlatformConfig(
name="官方文档站点",
platform_type=PlatformType.DOCS,
base_url="https://docs.example.com/",
content_limit=20000,
supports_structured=True,
ai_crawl_frequency="daily",
authority_score=0.90,
notes="自有文档站点,需做好SEO",
),
PlatformConfig(
name="Medium",
platform_type=PlatformType.BLOG,
base_url="https://medium.com/",
content_limit=8000,
supports_structured=False,
ai_crawl_frequency="daily",
authority_score=0.78,
notes="英文深度文章平台",
),
PlatformConfig(
name="微信公众号",
platform_type=PlatformType.BLOG,
base_url="https://mp.weixin.qq.com/",
content_limit=6000,
supports_structured=False,
ai_crawl_frequency="weekly",
authority_score=0.75,
notes="微信生态内搜索权重高",
),
PlatformConfig(
name="Stack Overflow",
platform_type=PlatformType.COMMUNITY,
base_url="https://stackoverflow.com/questions/",
content_limit=5000,
supports_structured=False,
ai_crawl_frequency="hourly",
authority_score=0.90,
notes="编程相关问题首选信源",
),
PlatformConfig(
name="CSDN",
platform_type=PlatformType.COMMUNITY,
base_url="https://blog.csdn.net/",
content_limit=8000,
supports_structured=False,
ai_crawl_frequency="daily",
authority_score=0.72,
notes="中文技术社区",
),
PlatformConfig(
name="掘金",
platform_type=PlatformType.COMMUNITY,
base_url="https://juejin.cn/post/",
content_limit=6000,
supports_structured=False,
ai_crawl_frequency="daily",
authority_score=0.70,
notes="前端/后端技术社区",
),
PlatformConfig(
name="Reuters/AP News",
platform_type=PlatformType.NEWS,
base_url="https://www.reuters.com/",
content原创声明:本文系作者授权腾讯云开发者社区发表,未经许可,不得转载。
如有侵权,请联系 cloudcommunity@tencent.com 删除。