✨ 研究情报看板

把密集日报转成适合人类阅读的研究雷达:先看主线,再看 Must Read,最后按方向筛选。

AI Research Radar - 2026-08-17

  • 研究画像:George Research Profile v2
  • 总结模式:单模型
  • 供应商:deepseek
  • 模型:deepseek-v4-flash
  • LLM 总结调用次数:7
  • 估算成本:RMB 0.0 / 1.0
  • 最近一次 LLM 错误:provider=deepseek; model=deepseek-v4-flash; base_url=https://api.deepseek.com; HTTP status=n/a; error=Could not parse JSON response:
  • 已禁用供应商:kimi
  • 原因:unauthorized

0. 每日概览

  • 最重要方向:上下文压缩 / 长上下文 / 记忆
  • 必读数量:0
  • 略读数量:8(Teaching LLMs to Update Beliefs for Efficient Long-Horizon Interaction;Adaptive Parallel Reasoning: The Next Paradigm in Efficient Inference Scaling;LycheeMemory V2: Efficient Long-Term Memory for LLM Agents via Semantic Segment-Level Consolidation;Mind the Context: Continual Learning of Socially Appropriate Robot Actions via Environmental-Social Disentanglement;Context-Matched Distillation: Teacher Causality for Autoregressive Video Distillation)
  • 关注数量:12(2026 BAIR Graduate Showcase;Identifying Interactions at Scale for LLMs;Spatial Memory Agent: Experience-Grounded Procedure Memory for Spatial Intelligence;Vero: Can AI Agents Build Formally Verified Software Repositories?;CoverPrune: Coverage-Driven Token Pruning for 3D VLMs via Optimal Transport)
  • 关键词:agent、framework、reasoning、agent memory、nlp、robotics、attention、berkeley.edu
  • 判断:今日主线:没有强制深读项,建议归档观察。

1. 核心研究方向

1.1 AI 系统 / HPC / 分布式训练与推理

必读

  • 无。

略读

  • 无。

关注

1.2 GPU 中心 I/O / 网络 / 存储

必读

  • 无。

略读

  • 无。

关注

1.3 AI 基础设施压缩 / 可靠性

必读

  • 无。

略读

  • 无。

关注

1.4 Agent 运行时 / RL 基础设施 / 调度

必读

  • 无。

略读

  • 无。

关注

1.5 具身智能 / VLA / 世界模型

必读

  • 无。

略读

1. Mind the Context: Continual Learning of Socially Appropriate Robot Actions via Environmental-Social Disentanglement
  • 阅读优先级:略读
  • 来源:arXiv AI/ML/NLP/Vision/Robotics(一手来源;角色=论文来源)
  • 发布时间:2026-08-13T16:33:10+00:00
  • 主方向:具身智能 / VLA / 世界模型
  • 次级标签:其他亮点、Novel Class Discovery / Open-World Learning / OOD / Continual Learning、AI 系统 / HPC / 分布式训练与推理、Agent 运行时 / RL 基础设施 / 调度
  • 依据层级:仅摘要
  • 评分:个人相关度=0.81,全局热度=0.47,可信度=1.00,证据强度=1.00,炒作风险=0.00,反馈=0.00
  • 项目相关性:skyfs=0.00、schedagent=0.00、verl_infrastructure=0.00、embodied_intelligence=0.22
  • 是什么:Mind the Context: Continual Learning of Socially Appropriate Robot Actions via Environmental-Social Disentanglement:研究论文,方向为“具身智能 / VLA / 世界模型”;主要线索:agent、continual learning、cs.RO、framework。
  • 问题:它关注“具身智能 / VLA / 世界模型”里的 agent、continual learning、cs.RO、framework 等问题。
  • 方法 / 贡献:摘要可确认它偏向评测或数据构建;具体任务定义、指标和样本规模需读原文确认。
  • 为什么对 George 重要:阅读优先级:略读 编辑优先级:0.74 今天快速扫读。 个人相关度:0.81,研究相关度:1.00。
  • 建议动作:快速扫读
  • 命中关键词:agent、continual learning、cs.RO、framework、github、nlp、robot、robotics

关注

2. 支撑性 AI 基础方向

上下文 / 记忆

  • Full-bandwidth transformer (关注;上下文压缩 / 长上下文 / 记忆;个人相关度=0.67;全局热度=0.43;炒作风险=0.00)

通用 Agent / 推理

强化学习

模型架构

多模态 / VLM / 计算机视觉

NLP

开放世界 / 持续学习

  • 无。

模型蒸馏

3. 跨方向连接

  • VLA inference latency ↔ GPU serving
  • robot rollout ↔ RL infrastructure
  • world model simulation ↔ HPC
  • KV cache ↔ storage hierarchy
  • gradient compression ↔ collective communication
  • agent workflow ↔ cluster scheduling
  • checkpoint ↔ GDS / distributed storage

4. Benchmark / 数据集 / 评测

Core Benchmarks for My Research

1. How Do VLMs Behave When Blind or Misled? Behavioral Evaluation of VLMs on Scientific Figures
  • 阅读层级:关注
  • 来源:arXiv AI/ML/NLP/Vision/Robotics
  • 证据来源:仅摘要
  • benchmark 评估什么能力:评估摘要中描述的任务能力;具体指标需打开原文确认。
  • 适合用于什么研究:适合用于评测协议、指标设计或负样本构造参考;是否纳入实验需看任务贴合度。
  • 可否作为实验基准:可以优先评估是否作为实验基准。
  • 建议行动:use_as_eval
2. MLLM-Routed Heterogeneous Ensembles for Robust Cross-Dataset Image Classification
  • 阅读层级:关注
  • 来源:arXiv AI/ML/NLP/Vision/Robotics
  • 证据来源:仅摘要
  • benchmark 评估什么能力:评估 agent 规划、执行或环境交互能力。
  • 适合用于什么研究:适合用于 agent evaluation / memory / long-horizon planning 相关实验。
  • 可否作为实验基准:可以优先评估是否作为实验基准。
  • 建议行动:use_as_eval
3. HumanTracker: Towards Comprehensive and Human-Aligned Motion Tracking Benchmark
  • 阅读层级:关注
  • 来源:arXiv AI/ML/NLP/Vision/Robotics
  • 证据来源:仅摘要
  • benchmark 评估什么能力:评估摘要中描述的任务能力;具体指标需打开原文确认。
  • 适合用于什么研究:适合用于评测协议、指标设计或负样本构造参考;是否纳入实验需看任务贴合度。
  • 可否作为实验基准:可以优先评估是否作为实验基准。
  • 建议行动:use_as_eval
4. TraVEL: Trajectory-Guided Video Embedding Learning for Driving-Video Retrieval
  • 阅读层级:关注
  • 来源:arXiv AI/ML/NLP/Vision/Robotics
  • 证据来源:仅摘要
  • benchmark 评估什么能力:评估摘要中描述的任务能力;具体指标需打开原文确认。
  • 适合用于什么研究:适合用于 agent evaluation / memory / long-horizon planning 相关实验。
  • 可否作为实验基准:可以优先评估是否作为实验基准。
  • 建议行动:use_as_eval
5. PlayWorld: Benchmarking World Models with Agent Players over Long-Horizon Objectives
  • 阅读层级:关注
  • 来源:arXiv AI/ML/NLP/Vision/Robotics
  • 证据来源:仅摘要
  • benchmark 评估什么能力:评估 agent 规划、执行或环境交互能力。
  • 适合用于什么研究:适合用于 agent evaluation / memory / long-horizon planning 相关实验。
  • 可否作为实验基准:可以优先评估是否作为实验基准。
  • 建议行动:use_as_eval

Interesting Benchmarks

1. H2R-Bench: Benchmarking Human-to-Robot Manipulation Video Generation in World Models
  • 阅读层级:关注
  • 来源:Hugging Face Daily Papers
  • 证据来源:仅摘要
  • benchmark 评估什么能力:评估摘要中描述的任务能力;具体指标需打开原文确认。
  • 适合用于什么研究:适合用于评测协议、指标设计或负样本构造参考;是否纳入实验需看任务贴合度。
  • 可否作为实验基准:暂不作为核心基准,先保存评测协议和指标设计。
  • 建议行动:save
2. Edit2TikZ: A Comprehensive and Challenging Benchmark for Scientific Figure Editing with TikZ
  • 阅读层级:关注
  • 来源:arXiv AI/ML/NLP/Vision/Robotics
  • 证据来源:仅摘要
  • benchmark 评估什么能力:评估摘要中描述的任务能力;具体指标需打开原文确认。
  • 适合用于什么研究:适合用于多模态泛化或跨域评测设计参考。
  • 可否作为实验基准:暂不作为核心基准,先保存评测协议和指标设计。
  • 建议行动:skim
3. Performance Evaluation of an Adaptive Quadrature and a Double Exponential Formula Using Arbitrary-Precision Floating-Point Arithmetic
  • 阅读层级:关注
  • 来源:arXiv Systems/HPC/GPU Data Path
  • 证据来源:仅摘要
  • benchmark 评估什么能力:评估摘要中描述的任务能力;具体指标需打开原文确认。
  • 适合用于什么研究:适合用于评测协议、指标设计或负样本构造参考;是否纳入实验需看任务贴合度。
  • 可否作为实验基准:暂不作为核心基准,先保存评测协议和指标设计。
  • 建议行动:skim
4. Evaluation of Clinically Steerable Retinal Image Generation from Foundation Model Latent Spaces
  • 阅读层级:关注
  • 来源:arXiv AI/ML/NLP/Vision/Robotics
  • 证据来源:仅摘要
  • benchmark 评估什么能力:评估摘要中描述的任务能力;具体指标需打开原文确认。
  • 适合用于什么研究:适合用于评测协议、指标设计或负样本构造参考;是否纳入实验需看任务贴合度。
  • 可否作为实验基准:暂不作为核心基准,先保存评测协议和指标设计。
  • 建议行动:save
5. Multi-UAV Tracking Evaluation Using 5G Uplink Signals on an O-RAN ISAC Simulation Testbed
  • 阅读层级:关注
  • 来源:arXiv Systems/HPC/GPU Data Path
  • 证据来源:仅摘要
  • benchmark 评估什么能力:评估摘要中描述的任务能力;具体指标需打开原文确认。
  • 适合用于什么研究:适合用于评测协议、指标设计或负样本构造参考;是否纳入实验需看任务贴合度。
  • 可否作为实验基准:暂不作为核心基准,先保存评测协议和指标设计。
  • 建议行动:skim

Other Benchmarks

  • 其余 4 个只进入附录标题列表:reports/appendix/2026-08-17-benchmarks.md

5. GitHub / 开源项目

New / Recently Active Projects

1. Paritok-official/paritok-4b-v1
  • 阅读优先级:研读代码
  • 来源:GitHub AI Research Projects(聚合来源;角色=代码可操作性来源)
  • 发布时间:2026-08-16T02:39:50+00:00
  • 主方向:GitHub / 开源项目推荐
  • 次级标签:上下文压缩 / 长上下文 / 记忆、Agent / 推理 / 推理时扩展 / 规划、AI 基础设施压缩 / 可靠性、Benchmark / 数据集 / 评测、工具库
  • 依据层级:仓库 README
  • 评分:个人相关度=0.69,全局热度=0.62,可信度=0.88,证据强度=0.69,炒作风险=0.00,反馈=0.00
  • 项目相关性:skyfs=0.00、schedagent=0.00、verl_infrastructure=0.00、embodied_intelligence=0.00
  • 是什么:Paritok-official/paritok-4b-v1:开源项目,方向为“GitHub / 开源项目推荐”;主要线索:agent、agentic、compression、context window。
  • 问题:它关注“GitHub / 开源项目推荐”里的 agent、agentic、compression、context window 等问题。
  • 方法 / 贡献:这是代码仓库条目;优先检查 README、示例、许可证和是否有可复现实验入口。
  • 为什么对 George 重要:阅读优先级:研读代码 编辑优先级:0.29 按 GitHub 项目动作处理。 个人相关度:0.69,研究相关度:0.69。
  • 建议动作:研读代码
  • 命中关键词:agent、agentic、compression、context window、evaluation、github、github.com、open-source
2. NousResearch/hermes-agent
  • 阅读优先级:克隆运行
  • 来源:GitHub AI Research Projects(聚合来源;角色=代码可操作性来源)
  • 发布时间:2026-08-16T22:43:12+00:00
  • 主方向:GitHub / 开源项目推荐
  • 次级标签:AI 系统 / HPC / 分布式训练与推理、Agent 运行时 / RL 基础设施 / 调度、工具库
  • 依据层级:仓库 README
  • 评分:个人相关度=0.81,全局热度=0.62,可信度=0.89,证据强度=0.69,炒作风险=0.00,反馈=0.00
  • 项目相关性:skyfs=0.00、schedagent=0.00、verl_infrastructure=0.00、embodied_intelligence=0.00
  • 是什么:NousResearch/hermes-agent:开源项目,方向为“GitHub / 开源项目推荐”;主要线索:GPU cluster、agent、cluster、github。
  • 问题:它关注“GitHub / 开源项目推荐”里的 GPU cluster、agent、cluster、github 等问题。
  • 方法 / 贡献:这是代码仓库条目;优先检查 README、示例、许可证和是否有可复现实验入口。
  • 为什么对 George 重要:阅读优先级:克隆运行 编辑优先级:0.35 按 GitHub 项目动作处理。 个人相关度:0.81,研究相关度:0.95。
  • 建议动作:克隆运行
  • 命中关键词:GPU cluster、agent、cluster、github、github.com、gpu、open-source
3. bytedance/deer-flow
  • 阅读优先级:克隆运行
  • 来源:GitHub AI Research Projects(聚合来源;角色=代码可操作性来源)
  • 发布时间:2026-08-16T15:54:06+00:00
  • 主方向:GitHub / 开源项目推荐
  • 次级标签:Agent / 推理 / 推理时扩展 / 规划、Agent 运行时 / RL 基础设施 / 调度、工具库
  • 依据层级:仓库 README
  • 评分:个人相关度=0.81,全局热度=0.62,可信度=0.89,证据强度=0.69,炒作风险=0.00,反馈=0.00
  • 项目相关性:skyfs=0.00、schedagent=0.00、verl_infrastructure=0.00、embodied_intelligence=0.00
  • 是什么:bytedance/deer-flow:开源项目,方向为“GitHub / 开源项目推荐”;主要线索:agent、agentic、framework、github。
  • 问题:它关注“GitHub / 开源项目推荐”里的 agent、agentic、framework、github 等问题。
  • 方法 / 贡献:这是代码仓库条目;优先检查 README、示例、许可证和是否有可复现实验入口。
  • 为什么对 George 重要:阅读优先级:克隆运行 编辑优先级:0.35 按 GitHub 项目动作处理。 个人相关度:0.81,研究相关度:0.94。
  • 建议动作:克隆运行
  • 命中关键词:agent、agentic、framework、github、github.com、long-horizon、multi-agent、open-source

Paper-linked Repos

1. deepseek-ai/DeepSeek-OCR
  • 阅读优先级:研读代码
  • 来源:GitHub AI Research Projects(聚合来源;角色=代码可操作性来源)
  • 发布时间:2026-01-27T03:45:14+00:00
  • 主方向:GitHub / 开源项目推荐
  • 次级标签:Agent / 推理 / 推理时扩展 / 规划、AI 系统 / HPC / 分布式训练与推理、Benchmark / 数据集 / 评测、AI 基础设施压缩 / 可靠性、工具库
  • 依据层级:仓库 README
  • 评分:个人相关度=0.65,全局热度=0.45,可信度=0.89,证据强度=0.69,炒作风险=0.00,反馈=0.00
  • 项目相关性:skyfs=0.00、schedagent=0.00、verl_infrastructure=0.00、embodied_intelligence=0.00
  • 是什么:deepseek-ai/DeepSeek-OCR:开源项目,方向为“GitHub / 开源项目推荐”;主要线索:compression、environment、eval、github。
  • 问题:它关注“GitHub / 开源项目推荐”里的 compression、environment、eval、github 等问题。
  • 方法 / 贡献:这是代码仓库条目;优先检查 README、示例、许可证和是否有可复现实验入口。
  • 为什么对 George 重要:阅读优先级:研读代码 编辑优先级:0.11 按 GitHub 项目动作处理。 个人相关度:0.65,研究相关度:0.69。
  • 建议动作:研读代码
  • 命中关键词:compression、environment、eval、github、github.com、image、inference、open-source
2. rednote-machine-learning/RedKnot
  • 阅读优先级:研读代码
  • 来源:GitHub AI Research Projects(聚合来源;角色=代码可操作性来源)
  • 发布时间:2026-07-10T06:18:48+00:00
  • 主方向:GitHub / 开源项目推荐
  • 次级标签:AI 系统 / HPC / 分布式训练与推理、上下文压缩 / 长上下文 / 记忆、其他亮点、GPU 中心 I/O / 网络 / 存储、工具库
  • 依据层级:仓库 README
  • 评分:个人相关度=0.61,全局热度=0.36,可信度=0.88,证据强度=0.69,炒作风险=0.00,反馈=0.00
  • 项目相关性:skyfs=0.00、schedagent=0.00、verl_infrastructure=0.00、embodied_intelligence=0.00
  • 是什么:rednote-machine-learning/RedKnot:开源项目,方向为“GitHub / 开源项目推荐”;主要线索:alignment、attention、github、github.com。
  • 问题:它关注“GitHub / 开源项目推荐”里的 alignment、attention、github、github.com 等问题。
  • 方法 / 贡献:这是代码仓库条目;优先检查 README、示例、许可证和是否有可复现实验入口。
  • 为什么对 George 重要:阅读优先级:研读代码 编辑优先级:0.10 按 GitHub 项目动作处理。 个人相关度:0.61,研究相关度:0.68。
  • 建议动作:研读代码
  • 命中关键词:alignment、attention、github、github.com、inference、long-context、open-source、serving
3. microsoft/MInference
  • 阅读优先级:克隆运行
  • 来源:GitHub AI Research Projects(聚合来源;角色=代码可操作性来源)
  • 发布时间:2026-04-08T08:04:38+00:00
  • 主方向:GitHub / 开源项目推荐
  • 次级标签:上下文压缩 / 长上下文 / 记忆、模型架构、AI 系统 / HPC / 分布式训练与推理、其他亮点、工具库
  • 依据层级:仓库 README
  • 评分:个人相关度=0.64,全局热度=0.48,可信度=0.88,证据强度=0.69,炒作风险=0.00,反馈=0.00
  • 项目相关性:skyfs=0.00、schedagent=0.00、verl_infrastructure=0.00、embodied_intelligence=0.00
  • 是什么:microsoft/MInference:开源项目,方向为“GitHub / 开源项目推荐”;主要线索:attention、github、github.com、inference。
  • 问题:它关注“GitHub / 开源项目推荐”里的 attention、github、github.com、inference 等问题。
  • 方法 / 贡献:这是代码仓库条目;优先检查 README、示例、许可证和是否有可复现实验入口。
  • 为什么对 George 重要:阅读优先级:克隆运行 编辑优先级:0.13 按 GitHub 项目动作处理。 个人相关度:0.64,研究相关度:0.65。
  • 建议动作:克隆运行
  • 命中关键词:attention、github、github.com、inference、long-context、open-source、release、sparse attention

Evergreen Toolkits

  • 今日无需要重复推荐的常青工具库。

6. 学者雷达

  • Jeff Dean: focus=ai_systems_hpc, distributed_systems, machine_learning_systems; last_verified=2026-07-18
  • Richard Sutton: focus=rl, agent_rl_infrastructure; last_verified=2026-07-18
  • Torsten Hoefler: focus=ai_systems_hpc, gpu_data_path_storage, compression_reliability; last_verified=2026-07-18
  • Pieter Abbeel: focus=embodied_world_models, rl; last_verified=2026-07-18
  • Shunyu Yao: focus=agent_rl_infrastructure, agents; last_verified=2026-07-18
  • 孙凝晖: focus=ai_systems_hpc, hpc; last_verified=2026-07-18
  • 赵海睿: focus=agent_rl_infrastructure, ai_systems_hpc; last_verified=2026-07-18

7. 高校 / 实验室雷达

8. 公司研究雷达

  • Stanford University: focus=ai, systems, robotics; last_verified=unverified
  • MIT: focus=ai_systems_hpc, robotics; last_verified=unverified
  • UC Berkeley: focus=systems, ai, robotics; last_verified=unverified
  • Carnegie Mellon University: focus=systems, robotics, ai; last_verified=unverified
  • Tsinghua University: focus=ai_systems_hpc, ai; last_verified=unverified
  • Institute of Computing Technology, CAS: focus=ai_systems_hpc, distributed_systems; last_verified=unverified
  • NVIDIA Research: focus=gpu_data_path_storage, ai_systems_hpc, embodied_world_models; last_verified=unverified
  • Google DeepMind: focus=ai, embodied_world_models, rl; last_verified=unverified

9. 学会 / 奖项 / Fellow / 领导层

学会

  • ACM: focus=computer_science, ai_systems_hpc; last_verified=2026-07-18
  • IEEE: focus=computer_science, electrical_engineering; last_verified=2026-07-18
  • IEEE Computer Society: focus=ai_systems_hpc, gpu_data_path_storage; last_verified=2026-07-18
  • AAAI: focus=ai, agent_rl_infrastructure; last_verified=2026-07-18
  • CCF: focus=computer_science; last_verified=2026-07-18
  • USENIX: focus=systems, security, storage; last_verified=2026-07-18
  • SIAM: focus=hpc, scientific_computing; last_verified=2026-07-18
  • ACL: focus=nlp; last_verified=2026-07-18

奖项与代表论文

  • 今日无高相关顶会精选。

10. 重要会议与期刊论文

  • NeurIPS: focus=not specified; last_verified=2026-07-18
  • ICML: focus=not specified; last_verified=2026-07-18
  • ICLR: focus=not specified; last_verified=2026-07-18
  • SOSP: focus=not specified; last_verified=2026-07-18
  • OSDI: focus=not specified; last_verified=2026-07-18
  • FAST: focus=not specified; last_verified=2026-07-18
  • SC: focus=not specified; last_verified=2026-07-18
  • SIGCOMM: focus=not specified; last_verified=2026-07-18

11. 常青经典

1. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks(2020)

  • 作者:Patrick Lewis、Ethan Perez、Aleksandra Piktus、Fabio Petroni、Vladimir Karpukhin、Naman Goyal、Heinrich Kuttler、Mike Lewis 等
  • topic_tags:context_compression、long_context、nlp
  • 关联方向:Context Compression / Long Context / Memory、NLP
  • 为什么经典:RAG 把参数记忆和外部检索连接起来,是理解今天检索增强、agent memory、记忆有效性和上下文预算取舍的关键参照。
  • 今日新论文继承了什么问题:今天的相关条目 延续了经典工作里的核心问题:有限上下文、外部记忆与状态复用如何支撑更长程的推理。
  • 它挑战了什么经典假设:它挑战的是静态检索、固定窗口或只读记忆的假设,转向会随新证据更新的工作记忆和缓存管理。
  • 它推进到什么新场景:新场景从语言建模推进到 agent memory、动态 workflow 和长上下文服务系统。
  • 预备知识:了解 seq2seq、dense retrieval 和生成式问答。

12. 反馈感知推荐

  • No explicit feedback signal yet; using cold-start research profile.

13. 来源健康状态

  • OpenReview:错误(0 条) - 返回内容为空或不是合法 JSON: line 1 column 1 (char 0)
  • GitHub AI Research Projects:time budget exhausted(23 条) - 时间预算已耗尽 after 23 items
  • MIT CSAIL News:超时(0 条) - timeout after 25s
  • The Batch by DeepLearning.AI:错误(0 条) - 403 Client Error: Forbidden for url: https://www.deeplearning.ai/the-batch

14. 采集说明

  • 生成时间:2026-08-16T22:51:31.297616+00:00
  • 来源数量:31
  • 原始条目数:692
  • 去重后条目数:563
  • API 请求总数:7
  • 各供应商 API 请求数:deepseek:6, kimi:1
  • 缓存命中:1
  • 缓存未命中:5
  • Benchmark 附录:reports/appendix/2026-08-17-benchmarks.md
  • 报告路径:reports/daily/2026/08/2026-08-17.md
  • 上一份报告链接:reports/daily/2026/08/2026-08-16.md