《通用人工智能行为准则》解读及其对人工智能大模型归档的启示
作者贡献声明/Author contributions:
王孟涵:论文撰写及修改;
吴志杰:论文指导及修改;
张 静:论文指导及修改。
|
王孟涵,硕士研究生; |
|
吴志杰,副研究馆员,博士。 |
收稿日期: 2026-04-12
网络出版日期: 2026-08-25
The General-Purpose AI Code of Practice:Interpretation and Archival Implications for AI Models
|
Wang Menghan, Master Candidate; |
|
Wu Zhijie, Associate Research Librarian, PhD. |
Received date: 2026-04-12
Online published: 2026-08-25
[目的/意义] 人工智能大模型的开发与应用具有分级监管风险、全生命周期动态治理等特殊要求,使得其归档有着不同于一般信息管理系统或平台归档的特殊性。对AI大模型归档研究,有助于适应当前人工智能发展形势,丰富档案管理工作。[方法/过程] 以《通用人工智能行为准则》为研究对象,分别运用BERTopic模型进行主题识别以梳理《准则》中关于AI大模型开发与应用的要求与主要风险,运用DeepSeek V4模型进行内容抽取以归纳面向风险治理的举措及其形成的文件材料。在此基础上,从强调分级监管AI大模型风险、实施全生命周期动态治理及注重完善治理架构3个方面系统分析与解读《准则》,并探讨档案部门作为后端管理主体应如何主动应对上述要求与风险。[结果/结论] 基于上述分析,本文提出档案部门可从实施与风险治理相适应的分级归档、明确AI大模型全生命周期的归档要求以及完善面向风险治理的文件材料归档范围3个方面着手,为我国开展人工智能大模型归档工作提供参考。
关键词: 《通用人工智能行为准则》; 人工智能大模型归档; 风险治理; 全生命周期归档
王孟涵 , 吴志杰 , 张静 . 《通用人工智能行为准则》解读及其对人工智能大模型归档的启示[J]. 知识管理论坛, 2026 , 11(4) : 345 -351 . DOI: 10.13266/j.issn.2095-5472.2026.029
[Purpose/Significance] The development and application of AI models are subject to specific regulatory requirements covering tiered risk regulation and dynamic full-lifecycle governance, resulting in archiving characteristics distinct from those of conventional information management systems or platforms. Research on the archiving of AI models can keep pace with the evolving landscape of artificial intelligence and enrich both the theory and practice of archival management. [Method/Process] This paper took the General-Purpose AI Code of Practice as the research object. It applied the BERTopic model for topic identification to outline the requirements and major risks concerning the development and application of AI models in the Code of Practice, and employed the DeepSeek V4 model for content extraction to summarize the documents and materials formed for risk governance. Based on this, the paper systematically analyzed and interpreted the Code of Practice from three dimensions: emphasizing tiered risk regulation for AI models, implementing dynamic full-lifecycle governance, and discussed how archival institutions, as back-end management subjects, should actively respond to the above requirements and risks. [Result/Conclusion] Based on the above analysis, this paper proposes that archival institutions can advance the archiving of AI models from three dimensions: implementing tiered archiving matching risk governance, clarifying full-lifecycle archiving specifications, and optimizing the scope of records collection oriented toward risk management, to provide references for the domestic practice of AI models archiving in China.
表1 BERTopic模型识别的主题Table 1 Topics recognized by BERTopic model |
| 序号 | 主题号 | 关键词 |
|---|---|---|
| 1 | Topic0 | copyright related、related rights、law copyright、union law、subject matter |
| 2 | Topic1 | applicable national、office applicable、involvement model、signatories aware、relevant information |
| 3 | Topic2 | model parameters、reduction risk、security mitigations、mitigation objectives、insider threats |
| 4 | Topic3 | ai act、model documentation、obligations ai、downstream providers、providers ai |
| 5 | Topic4 | model report、model intended、updated model、provide ai、signatories provide |
| 6 | Topic5 | management body、body supervisory、working group、managing systemic、mitigation processes |
| 7 | Topic6 | signatories recognise、model evaluation、model evaluations、external evaluators、model version |
| 8 | Topic7 | risk identification、risks measure、types risks、understanding systemic、identification process |
| 9 | Topic8 | risk acceptance、specified measure、risk tiers、specified commitment、measure signatories |
表2 《准则》内容抽取Table 2 Content extraction from the Code of Practice |
| 序号 | 要求 | 举措 | 形成文件材料 |
|---|---|---|---|
| 1 | 透明度 | 提供并定期更新模型文档表 | 模型文档表、模型文档表更新日志 |
| 2 | 公开联系方式 | 模型提供者联系信息 | |
| 3 | 确保文档信息进行质量与完整性控制 | 信息质量与完整性控制记录 | |
| 4 | 版权合规 | 实施并持续更新版权政策 | 版本政策文档 |
| 5 | 识别和排除大规模侵权网站 | 网站超链接列表 | |
| 6 | 网络爬虫需要遵守权利保留声明 | 机器可读权利保留协议清单 | |
| 7 | 公开爬虫信息和遵守权利声明的技术规则,并通知权利人相关变化 | 网络爬虫信息及权利人通知记录 | |
| 8 | 实施技术保护,防止模型侵权性复制训练内容 | 侵权防护技术保障举措文档 | |
| 9 | 指定版权持有人电子联系点并公开信息 | 版权联系点信息 | |
| 10 | 建立版权投诉机制,公开渠道,及时处理 | 版权投诉机制文档及处理记录 | |
| 11 | 安全与风险治理 | 创建、实施并更新安全与风险治理框架 | 安全与风险治理框架、框架更新日志 |
| 12 | 在模型投放市场前,编制安全报告并提交至AI办公室 | 模型安全报告 | |
| 13 | 模型上市前进行外部安全审查、红队测试 | 安全审查报告、红队测试记录 | |
| 14 | 实施沙箱隔离、代码执行隔离、训练数据完整性检查 | 沙箱配置文档、隔离策略、数据完整性验证记录 | |
| 15 | 模型上市后持续监测收集风险评估信息 | 上市后监测数据 | |
| 16 | 对已识别的系统性风险进行风险识别、建模和估算 | 系统性风险清单、风险建模结果、风险估算结果、风险情景清单 | |
| 17 | 判定系统性风险是否可接受 | 风险接受度判定文档及决策记录 | |
| 18 | 过滤和清理训练数据、监控输入/输出等风险缓解举措 | 风险缓解举措实施记录文档 | |
| 19 | 持续跟踪和报告严重事件 | 严重事件跟踪、记录和报告相关信息 | |
| 20 | 明确组织各层级的系统性风险管理职责 | 系统性风险责任分配文档 | |
| 21 | 允许挑战风险、设置激励举措、举报人保护 | 员工调查、举报人保护政策 | |
| 22 | 配备充足人力、财力等资源 | 资源分配记录 |
表3 面向风险治理的AI大模型文件材料归档内容Table 3 Archiving scope of AI models documents and materials for risk |
| 序号 | 风险等级 | 相关文件材料内容 |
|---|---|---|
| 1 | 普通风险 | 模型文档表、模型文档表更新日志、模型提供者联系信息、信息质量与完整性控制记录、版本政策文档、网站超链接列表、机器可读权利保留协议清单、网络爬虫信息及权利人通知记录、侵权防护技术保障举措文档、版权联系点信息、版权投诉机制文档及处理记录等 |
| 2 | 系统性风险 | 安全与风险治理框架、框架更新日志、模型安全报告、安全审查报告、红队测试记录、沙箱配置文档、隔离策略、数据完整性验证记录、上市后监测数据、系统性风险清单、风险建模结果、风险估算结果、风险情景清单、风险接受度判定文档及决策记录、风险缓解举措实施记录文档、严重事件跟踪、记录和报告相关信息、系统性风险责任分配文档、员工调查、举报人保护政策、资源分配记录等 |
| [1] |
贺谭涛.档案学视角下的人工智能文档: 范畴界定、管理挑战与探索方向[J]. 档案学通讯, 2026(4): 39-47.
|
| [2] |
连志英, 蒋玲.面向人工智能治理的人工智能文档的作用及其构成研究[J]. 图书情报工作, 2026, 70(9): 148-156.
|
| [3] |
张茜雅, 刘越男.并行数据: 人工智能应用背景下档案管理的新议题[J]. 浙江档案, 2025(7): 14-19.
|
| [4] |
王阿陶.“算法档案”的概念内涵、治理理念与实现机制[J]. 档案学研究, 2026(2): 30-39.
|
| [5] |
连志英, 苏立.档案学视角下算法来源的内涵及其要素构成[J]. 档案学通讯, 2025(3): 29-37.
|
| [6] |
裴佳杰, 张斌.AIGC归档的逻辑理路与冲突解构[J]. 情报科学, 2025, 43(8): 100-108.
|
| [7] |
王玉珏, 樊静雅, 温翰英.人工智能生成合成内容的可信存档策略研究——基于对电子档案“四性”的思考[J]. 北京档案, 2025(4): 22-29.
|
| [8] |
|
| [9] |
蔡盈芳.科学界定科研文件材料的归档范围并准确划分科研档案保管期限[J]. 中国档案, 2021(5): 44-45.
|
| [10] |
深度求索.DeepSeek-V4预览版: 迈入百万上下文普惠时代[EB/OL]. [2026-05-30].
DeepSeek. DeepSeek-V4 preview: entering the era of accessible million-token context[EB/OL]. [2026-05-30].
|
| [11] |
徐拥军, 闫静.中国特色档案学的基本范畴与核心命题[J]. 中国图书馆学报, 2024, 50(3): 30-46.
|
| [12] |
田泽懿, 张靖琦.基于生命周期的大模型安全审视: 风险梳理与防范机制[J]. 工业信息安全, 2025(4): 20-28.
|
/
| 〈 |
|
〉 |