<?xml version="1.0" encoding="utf-8" ?><rss version="2.0"><channel><title><![CDATA[志在山顶的人，不会贪念山腰的风景]]></title><description><![CDATA[]]></description><link>https://blog.csdn.net/m0_61676839</link><language>zh-cn</language><generator>https://blog.csdn.net/</generator><copyright><![CDATA[Copyright &copy; m0_61676839]]></copyright><item><title><![CDATA[解决 TensorBoard 启动报错：ModuleNotFoundError: No module named ‘pkg_resources‘]]></title><link>https://blog.csdn.net/m0_61676839/article/details/161428069</link><guid>https://blog.csdn.net/m0_61676839/article/details/161428069</guid><author>m0_61676839</author><pubDate>Tue, 26 May 2026 17:44:53 +0800</pubDate><description><![CDATA[遇到缺失？别慌，大概率是setuptools版本太高惹的祸。]]></description><category></category></item><item><title><![CDATA[【2026 CVPR】ProactiveMobile: A Comprehensive Benchmark for Boosting Proactive Intelligence on Mobile]]></title><link>https://blog.csdn.net/m0_61676839/article/details/160666812</link><guid>https://blog.csdn.net/m0_61676839/article/details/160666812</guid><author>m0_61676839</author><pubDate>Thu, 30 Apr 2026 19:49:34 +0800</pubDate><description><![CDATA[TIkFkk1nPredictUDWBT{(Ik​Fk​k1n​PredictUDWB其中UDWBU, D, W, BUDWB分别代表上述四类上下文信号，IkI_kIk​是意图，FkF_kFk​是对应的函数序列。对和MiMo-VL-7B进行了全参数微调。训练模型同时输出“自然语言建议”和“执行函数序列”，研究发现这种联合输出（逻辑支架）能显著提升成功率并降低错误触发率。]]></description><category></category></item><item><title><![CDATA[MySQL 与向量数据库的核心区别：从结构化数据到语义搜索]]></title><link>https://blog.csdn.net/m0_61676839/article/details/160613017</link><guid>https://blog.csdn.net/m0_61676839/article/details/160613017</guid><author>m0_61676839</author><pubDate>Wed, 29 Apr 2026 09:42:05 +0800</pubDate><description><![CDATA[在数据技术不断演进的今天，传统数据库已经无法完全满足人工智能时代的需求。尤其是在大模型（LLM）和语义搜索兴起之后，一类新的数据库——向量数据库，逐渐成为热门选择。那么，经典的 MySQL 与向量数据库到底有什么本质区别？它们是否会相互取代？]]></description><category></category></item><item><title><![CDATA[【2025 CVPR】EMOE: Modality-Specific Enhanced Dynamic Emotion Experts]]></title><link>https://blog.csdn.net/m0_61676839/article/details/160600280</link><guid>https://blog.csdn.net/m0_61676839/article/details/160600280</guid><author>m0_61676839</author><pubDate>Tue, 28 Apr 2026 20:17:19 +0800</pubDate><description><![CDATA[摘要： EMOE提出了一种新颖的多模态情感识别方法，针对现有融合方法的两大核心问题：模态平衡困境和模态特殊性丧失。通过专家混合机制(MoME)实现样本级动态融合，并引入单模态蒸馏(UD)保留各模态的独立预测能力。在CMU-MOSI/MOSEI等基准测试中，EMOE显著优于现有方法，最高提升1.5%准确率。其创新性体现在：1）自适应权重路由网络，2）知识蒸馏保持模态特异性，3）良好的跨任务泛化性。代码已开源，为多模态学习提供了新思路。]]></description><category></category></item><item><title><![CDATA[从 Chain-of-Thought 到 Graph of Thoughts：LLM 推理范式的演进]]></title><link>https://blog.csdn.net/m0_61676839/article/details/160566252</link><guid>https://blog.csdn.net/m0_61676839/article/details/160566252</guid><author>m0_61676839</author><pubDate>Mon, 27 Apr 2026 20:36:52 +0800</pubDate><description><![CDATA[CoT 让模型“会解释”，ToT 让模型“会探索”，GoT 让模型“会思考”。]]></description><category></category></item><item><title><![CDATA[Prompt Engineering、Context Engineering、Harness Engineering三者之间的区别]]></title><link>https://blog.csdn.net/m0_61676839/article/details/160565739</link><guid>https://blog.csdn.net/m0_61676839/article/details/160565739</guid><author>m0_61676839</author><pubDate>Mon, 27 Apr 2026 19:27:19 +0800</pubDate><description><![CDATA[👉 重点：你给模型加了额外上下文（RAG / memory / docs）👉 重点：不只是一次调用，而是一个流程（带搜索 + 校验 + 兜底）👉 重点：你写了什么 prompt。]]></description><category></category></item><item><title><![CDATA[服务器 CUDA版本升级指南]]></title><link>https://blog.csdn.net/m0_61676839/article/details/160286967</link><guid>https://blog.csdn.net/m0_61676839/article/details/160286967</guid><author>m0_61676839</author><pubDate>Sat, 18 Apr 2026 21:54:58 +0800</pubDate><description><![CDATA[检查软链接：如果您的系统中安装了多个版本的CUDA，可能需要更新软链接/usr/local/cuda指向新版本的CUDA。反应的是显卡驱动 (Driver) 能够支持的 CUDA 最大版本，它决定了你能运行多高版本的 Toolkit，但并不强制要求你的项目环境必须使用这个最高版本。查到的 CUDA 版本（驱动支持的最高版本），并不等同于当前环境实际调用的 CUDA Toolkit 版本（这样，/usr/local/cuda就会指向CUDA 12.4的安装目录。安装目标 CUDA 版本。]]></description><category></category></item><item><title><![CDATA[flash-attn安装指南]]></title><link>https://blog.csdn.net/m0_61676839/article/details/160286376</link><guid>https://blog.csdn.net/m0_61676839/article/details/160286376</guid><author>m0_61676839</author><pubDate>Sat, 18 Apr 2026 20:59:57 +0800</pubDate><description><![CDATA[的作用是在不牺牲模型精度的前提下，让注意力机制（Attention Mechanism）跑得更快、更省内存。核心痛点：传统 Attention 需要生成巨大的N×NN \times NN×N矩阵，导致 GPU 显存频繁在“慢速大内存 (HBM)”与“快速小内存 (SRAM)”间搬运数据。GPU 算力极强，但因为忙着搬数据，大部分时间在“空转”。核心方案将矩阵切成小块，在高速 SRAM 中完成计算，无需回写到慢速内存。实时处理数值，省去保存完整中间矩阵的步骤。]]></description><category></category></item><item><title><![CDATA[【2026 ICLR】MME-Emotion: A Holistic Evaluation Benchmark for Emotional Intelligence in MLLM]]></title><link>https://blog.csdn.net/m0_61676839/article/details/160146479</link><guid>https://blog.csdn.net/m0_61676839/article/details/160146479</guid><author>m0_61676839</author><pubDate>Tue, 14 Apr 2026 14:17:23 +0800</pubDate><description><![CDATA[一个专门用于评估多模态大语言模型（MLLMs）情感智能的综合基准测试。]]></description><category></category></item><item><title><![CDATA[Git使用]]></title><link>https://blog.csdn.net/m0_61676839/article/details/160031499</link><guid>https://blog.csdn.net/m0_61676839/article/details/160031499</guid><author>m0_61676839</author><pubDate>Fri, 10 Apr 2026 20:20:33 +0800</pubDate><description><![CDATA[（保留暂存区）或默认不写（保留工作区）。中指定的文件，不会出现在。在文件中添加需要忽略的内容。之后，Git 会自动忽略。先在本地创建对应分支。（彻底回退，谨慎用）]]></description><category></category></item><item><title><![CDATA[Vue 渲染 Markdown 完整指南]]></title><link>https://blog.csdn.net/m0_61676839/article/details/160021999</link><guid>https://blog.csdn.net/m0_61676839/article/details/160021999</guid><author>m0_61676839</author><pubDate>Fri, 10 Apr 2026 14:48:19 +0800</pubDate><description><![CDATA[在开发 AI 对话应用、技术文档站点或博客系统时，经常需要在 Vue 中渲染 Markdown 内容。本文总结了 3 种主流方案，从轻量级到专业化，帮你快速选型并落地。核心修正：此方案主要用于纯前端渲染。如果是 Nuxt.js 等 SSR 项目，请改用  或 。

🚀 方案二：专业解析 (markdown-it)

主要特点：

💬 方案三：AI 流式对话 (Vue Stream Markdown)

逻辑优化：
2. 封装流式组件


🛡️ 关键注意事项 (SSR 与 安全)


SSR (Nuxt]]></description><category></category></item><item><title><![CDATA[【2026 AAAI】Causal-ERC: A Multimodal Framework with Causal Prompting for Emotion Recognition in Conve]]></title><link>https://blog.csdn.net/m0_61676839/article/details/159964278</link><guid>https://blog.csdn.net/m0_61676839/article/details/159964278</guid><author>m0_61676839</author><pubDate>Wed, 08 Apr 2026 19:26:15 +0800</pubDate><description><![CDATA[IEMOCAP：包含演员表演的剧本对话，情感标签包括快乐、愤怒、中性、悲伤、兴奋和沮丧。MELD：源自美剧《老友记》的多方对话数据集，包含中性、惊喜、恐惧、悲伤、快乐、厌恶和愤怒等标签。模态组成：每个对话切片（Utterance）均包含文本（Textual）音频（Acoustic）和视觉（Visual）三种特征。]]></description><category></category></item><item><title><![CDATA[【2026 AAAI】Beyond Counting: Evaluating Abstract and Emotional Reasoning in Vision-Language Models]]></title><link>https://blog.csdn.net/m0_61676839/article/details/159963715</link><guid>https://blog.csdn.net/m0_61676839/article/details/159963715</guid><author>m0_61676839</author><pubDate>Wed, 08 Apr 2026 19:00:16 +0800</pubDate><description><![CDATA[由于 EmojiGrid 的图像是由表情符号（Emoji）组成的，这些符号具有“超越语言”的通用性（例如，全世界的人都能看懂“😢”代表悲伤）。：设计上兼顾了各类任务的比例，防止模型通过过拟合某种任务来刷分。例如感知类占40.43%，关系类占32.12%，抽象类占27.44%。最近的视觉语言模型（如 Gemini 2.5 Pro, GPT-o1/o4-mini, GLM-4V-Thinking）都引入了。: 问题的 Token 长度分布呈现长尾模式，并通过 Emoji 占格子的比例来衡量视觉干扰度。]]></description><category></category></item><item><title><![CDATA[【2026 arXiv】EVA: Efficient Reinforcement Learning for End-to-End Video Agent]]></title><link>https://blog.csdn.net/m0_61676839/article/details/159862757</link><guid>https://blog.csdn.net/m0_61676839/article/details/159862757</guid><author>m0_61676839</author><pubDate>Sun, 05 Apr 2026 21:30:07 +0800</pubDate><description><![CDATA[研究在六个主要的视频理解基准测试上进行了评估，包括 LSDBench、VideoMME 和 LongVideoBench 等。采样效率：在采样困境基准（LSDBench）上，EVA 仅使用极小的视觉 Token 就达到了51.8%的准确率，优于使用大量 Token 的基线模型（如 Qwen2.5-VL 256 帧需要 166.4K Token）。长视频理解：在 VideoMME 等长视频任务中，EVA 展现出一致的领先优势，证明了其自适应关注关键片段的能力。推理能力。]]></description><category></category></item><item><title><![CDATA[【2026 CVPR】Asking like Socrates: Socrates helps VLMs understand remote sensing images]]></title><link>https://blog.csdn.net/m0_61676839/article/details/159861741</link><guid>https://blog.csdn.net/m0_61676839/article/details/159861741</guid><author>m0_61676839</author><pubDate>Sun, 05 Apr 2026 20:50:58 +0800</pubDate><description><![CDATA[研究旨在解决视觉语言模型（VLM）在处理遥感图像时的“虚假推理”问题。]]></description><category></category></item><item><title><![CDATA[【2026 arXiv】OmniRAG-Agent: Agentic Omnimodal Reasoning for Low-Resource Long Audio-Video Question A]]></title><link>https://blog.csdn.net/m0_61676839/article/details/158965893</link><guid>https://blog.csdn.net/m0_61676839/article/details/158965893</guid><author>m0_61676839</author><pubDate>Thu, 12 Mar 2026 21:09:59 +0800</pubDate><description><![CDATA[摘要： 针对低资源环境下全模态大语言模型（OmniLLMs）处理长音视频问答的挑战，本文提出 OmniRAG-Agent 框架，结合检索增强生成（RAG）、多轮代理规划和强化学习（RL）。通过构建图像-音频双库实现细粒度检索，代理机制自主调用工具跨轮次整合证据，并采用组相对策略优化（GRPO）联合提升工具使用与答案生成能力。实验表明，该方法在多个基准测试中显著优于基线模型，尤其在细粒度检索、逻辑推理等任务上表现突出，且具备良好的泛化性和骨干网络迁移性。核心贡献包括验证低预算检索可行性、多步代理规划的有效性，]]></description><category></category></item><item><title><![CDATA[【2025 arXiv】Prompt-R1: Collaborative Automatic Prompting Framework via End-to-end Reinforcement Lear]]></title><link>https://blog.csdn.net/m0_61676839/article/details/158976964</link><guid>https://blog.csdn.net/m0_61676839/article/details/158976964</guid><author>m0_61676839</author><pubDate>Thu, 12 Mar 2026 21:07:57 +0800</pubDate><description><![CDATA[多轮交互机制 (Multi-Turn Interaction)Agent 首先对问题进行初步思考（Think），生成初始提示词发送给大模型。Agent初始的提示词模板：大模型返回响应后，Agent 结合历史记录进行进一步推理，并调整下一轮提示词。这种往复过程持续到 Agent 认为已准备好输出最终答案。双重约束奖励 (Double-constrained Reward)格式奖励 (RfmtR_{fmt}Rfmt​：强制 Agent 遵循推理步骤，确保输出非空、可解析且格式正确。]]></description><category></category></item><item><title><![CDATA[【2026 CVPR】EMO-R3: Reflective Reinforcement Learning for Emotional Reasoning in MLLM]]></title><link>https://blog.csdn.net/m0_61676839/article/details/158892987</link><guid>https://blog.csdn.net/m0_61676839/article/details/158892987</guid><author>m0_61676839</author><pubDate>Tue, 10 Mar 2026 23:58:19 +0800</pubDate><description><![CDATA[本文提出EMO-R3框架，通过结构化情感思维(SET)和反射情感奖励(RER)增强多模态大语言模型的情感推理能力。SET引导模型分三阶段推理：情感触发点识别、人类情感反射和情感结论；RER则通过图文一致性和情感连贯性奖励实现自我评估。实验表明，该方法在EmoSet等数据集上优于现有技术，消融研究验证了各模块的有效性。主要贡献包括：1)结构化情感推理过程；2)反射式自我评估机制；3)在多个基准测试中的性能提升。该研究为提升AI情感理解能力提供了新思路。]]></description><category></category></item><item><title><![CDATA[【2026 arXiv】VoiceSculptor: Your Voice, Designed By You]]></title><link>https://blog.csdn.net/m0_61676839/article/details/157359559</link><guid>https://blog.csdn.net/m0_61676839/article/details/157359559</guid><author>m0_61676839</author><pubDate>Tue, 27 Jan 2026 20:36:14 +0800</pubDate><description><![CDATA[VoiceSculptor是一个开源的统一语音合成框架，通过自然语言指令实现精细化语音控制。其创新点包括：1）构建9,000小时多层次标注数据集，结合ASR、情感分析和韵律特征离散化；2）采用双阶段解耦架构（语音设计+克隆模块），基于LLaSA-3B和Cosy Voice2模型；3）引入思维链机制显式分解语音属性，并通过检索增强提升指令泛化能力。实验表明，该系统在中文指令基准测试中达到SOTA水平，消融实验验证了CoT和RAG的有效性。该工作填补了开源TTS系统在指令遵循能力上的技术空白。]]></description><category></category></item><item><title><![CDATA[【2026 AAAI】ClearAIR: A Human-Visual-Perception-Inspired All-in-One Image Restoration]]></title><link>https://blog.csdn.net/m0_61676839/article/details/157336378</link><guid>https://blog.csdn.net/m0_61676839/article/details/157336378</guid><author>m0_61676839</author><pubDate>Tue, 27 Jan 2026 20:35:20 +0800</pubDate><description><![CDATA[本文提出ClearAIR框架，这是一种受人类视觉感知启发的全能型图像修复方法。传统方法存在空间均匀性假设局限和细节丢失问题。ClearAIR通过四个阶段模拟人类视觉处理：1）MLLM-IQA模块进行全局质量评估；2）语义引导单元定位受损区域；3）任务识别器判断退化类型；4）内部线索重用机制恢复微观纹理。实验表明，ClearAIR在去噪、去雾等任务中性能显著优于现有方法，特别是在复合退化场景下PSNR提升0.62dB。该框架创新性地融合了多模态评估和自监督学习，实现了更精准、更自然的图像修复效果。]]></description><category></category></item></channel></rss>