<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>AI Native Foundation</title>
	<atom:link href="https://ainativefoundation.org/feed/" rel="self" type="application/rss+xml" />
	<link>https://ainativefoundation.org/</link>
	<description></description>
	<lastBuildDate>Mon, 21 Sep 2026 09:02:25 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.1</generator>

<image>
	<url>https://ainativefoundation.org/wp-content/uploads/2024/05/cropped-favicon-32x32.png</url>
	<title>AI Native Foundation</title>
	<link>https://ainativefoundation.org/</link>
	<width>32</width>
	<height>32</height>
</image> 
	<item>
		<title>Global AI Native Industry Insights &#8211; 20260921 &#8211;  xAI &#124; Anthropic &#124; OpenAI &#124; more</title>
		<link>https://ainativefoundation.org/global-ai-native-industry-insights-20260921-xai-anthropic-openai-more/</link>
		
		<dc:creator><![CDATA[AINF]]></dc:creator>
		<pubDate>Mon, 21 Sep 2026 09:02:25 +0000</pubDate>
				<category><![CDATA[Global Industry]]></category>
		<guid isPermaLink="false">https://ainativefoundation.org/?p=10998</guid>

					<description><![CDATA[<p>Today's digest highlights notable developments from xAI and Anthropic. The central theme revolves around advancements in AI capabilities and investment in evaluation frameworks. xAI introduces Grok Voice Transcribe 2.0, boasting top speech transcription accuracy, while Anthropic, in partnership with Accenture, plans to invest $1 billion in independent AI evaluation systems. Additionally, OpenAI brings multi-account support to its Codex plugins, enhancing user accessibility. Perplexity Computer adds app connections to its homepage composer, expanding its functionality and user interaction.</p>
<p>The post <a href="https://ainativefoundation.org/global-ai-native-industry-insights-20260921-xai-anthropic-openai-more/">Global AI Native Industry Insights &#8211; 20260921 &#8211;  xAI | Anthropic | OpenAI | more</a> appeared first on <a href="https://ainativefoundation.org">AI Native Foundation</a>.</p>
]]></description>
										<content:encoded><![CDATA[<p>Today&#8217;s digest highlights notable developments from xAI and Anthropic. The central theme revolves around advancements in AI capabilities and investment in evaluation frameworks. xAI introduces Grok Voice Transcribe 2.0, boasting top speech transcription accuracy, while Anthropic, in partnership with Accenture, plans to invest $1 billion in independent AI evaluation systems. Additionally, OpenAI brings multi-account support to its Codex plugins, enhancing user accessibility. Perplexity Computer adds app connections to its homepage composer, expanding its functionality and user interaction. Discover more in Today&#8217;s Global AI Native Industry Insights.</p>
<h3>1.  xAI Launches Grok Voice Transcribe 2.0, Claiming Top Speech Transcription Accuracy</h3>
<p>xAI announced Grok Voice Transcribe 2.0, a new version of its speech transcription model. The company states it is the world&#8217;s most accurate speech transcription model currently available. The announcement was made via xAI&#8217;s official account on September 18, 2026. No additional technical benchmarks or availability details were included in the announcement itself.<br />
Read more: <a href="https://x.ai/news/grok-voice-transcribe-2">https://x.ai/news/grok-voice-transcribe-2</a></p>
<p><video width="600" height="400" controls poster="https://cdn.ainative.foundation/image/20260921_01495d6c06dc44378d8e838c7b9d1fd7.jpg"><source src="https://cdn.ainative.foundation/video/20260921_gj_video_xai.mp4" type="video/mp4"></video><br />
Video Credit: NotebookLM</p>
<h3>2.  Anthropic and Accenture to Invest $1 Billion in Independent AI Evaluation</h3>
<p>Anthropic announced a partnership with Accenture to conduct independent evaluation of frontier AI systems. The collaboration follows Anthropic&#8217;s recent commitment to embed independent evaluators within the company. Both organizations plan to invest at least $1 billion combined over the next five years to build evaluation capacity in this area. The initiative aims to strengthen third-party assessment of advanced AI models.<br />
Read more: <a href="https://www.anthropic.com/news/accenture-embedded-evaluation">https://www.anthropic.com/news/accenture-embedded-evaluation</a></p>
<p><video width="600" height="400" controls poster="https://cdn.ainative.foundation/image/20260921_gj_img_anthropicai.jpg"><source src="https://cdn.ainative.foundation/video/20260921_gj_video_anthropic.mp4" type="video/mp4"></video><br />
Video Credit: NotebookLM</p>
<h3>3.  OpenAI Rolls Out Multi-Account Support Across Codex Plugins</h3>
<p>OpenAI&#8217;s developer team announced that multi-account support is now available across most Codex plugins. The update allows developers to switch between personal, work, and side-project accounts while building. This is a developer tool update aimed at improving workflow flexibility for users managing multiple accounts.<br />
Read more: <a href="https://x.com/i/web/status/2100980899655389392">https://x.com/i/web/status/2100980899655389392</a></p>
<p><video width="600" height="400" controls poster="https://cdn.ainative.foundation/image/20260921_gj_img_openai.jpg"><source src="https://cdn.ainative.foundation/video/20260921_gj_video_openai.mp4" type="video/mp4"></video><br />
Video Credit: NotebookLM</p>
<h3>4.  Perplexity Computer Adds App Connections to Homepage Composer</h3>
<p>Perplexity Computer has rolled out a new feature allowing users to connect apps directly from the homepage composer. The update is live now on the web and available to all Computer users. This expands the platform&#8217;s integration capabilities, enabling users to link external apps within their workflow directly from the main interface.<br />
Read more: <a href="https://x.com/i/web/status/2101147964307849718">https://x.com/i/web/status/2101147964307849718</a></p>
<p><video width="600" height="400" controls poster="https://cdn.ainative.foundation/image/20260921_871db123bfd74e52a36447a0aaf1c1bd.jpg"><source src="https://cdn.ainative.foundation/video/20260921_adad10fb46404f3eb461a82b21b6ebe1.mp4" type="video/mp4"></video><br />
Video Credit: @AskPerplexity on X</p>
<div style="width:100%;height:2px;background:#808080;margin:10px 0"></div>
<p>That&#8217;s all for today&#8217;s Global AI Native Industry Insights. Join us at <a href="https://member.ainativefoundation.org/">AI Native Foundation Membership Dashboard</a> for the latest insights on AI Native, or follow our linkedin account at <a href="https://www.linkedin.com/company/ainativefoundation/">AI Native Foundation</a> and our twitter account at <a href="https://x.com/AINativeF">AINativeF</a>.</p>
<p>The post <a href="https://ainativefoundation.org/global-ai-native-industry-insights-20260921-xai-anthropic-openai-more/">Global AI Native Industry Insights &#8211; 20260921 &#8211;  xAI | Anthropic | OpenAI | more</a> appeared first on <a href="https://ainativefoundation.org">AI Native Foundation</a>.</p>
]]></content:encoded>
					
		
		<enclosure url="https://cdn.ainative.foundation/video/20260921_gj_video_xai.mp4" length="6529848" type="video/mp4" />
<enclosure url="https://cdn.ainative.foundation/video/20260921_gj_video_anthropic.mp4" length="5341792" type="video/mp4" />
<enclosure url="https://cdn.ainative.foundation/video/20260921_gj_video_openai.mp4" length="4615177" type="video/mp4" />
<enclosure url="https://cdn.ainative.foundation/video/20260921_adad10fb46404f3eb461a82b21b6ebe1.mp4" length="2421561" type="video/mp4" />

			</item>
		<item>
		<title>China AI Native Industry Insights &#8211; 20260920 &#8211;  Z.ai &#124; MiniMax &#124; Alibaba &#124; more</title>
		<link>https://ainativefoundation.org/china-ai-native-industry-insights-20260920-z-ai-minimax-alibaba-more/</link>
		
		<dc:creator><![CDATA[AINF]]></dc:creator>
		<pubDate>Sun, 20 Sep 2026 08:41:52 +0000</pubDate>
				<category><![CDATA[China Industry]]></category>
		<guid isPermaLink="false">https://ainativefoundation.org/?p=10995</guid>

					<description><![CDATA[<p>Today's digest highlights some of the most compelling advancements from Z.ai, MiniMax, and Alibaba. These developments underscore a broader theme of enhanced efficiency and functionality in AI tools. Z.ai has significantly increased throughput with its GLM-5.3 system supporting the GLM-5.3-Flash, while StepFun introduces a preview of its Step 5 model, boasting a sparse MoE flagship with a striking 600 billion parameters and a 1 million token context window. Furthermore, Alibaba's introduction of Qwen3.8-LiveTranslate advances real-time speech translation capabilities.</p>
<p>The post <a href="https://ainativefoundation.org/china-ai-native-industry-insights-20260920-z-ai-minimax-alibaba-more/">China AI Native Industry Insights &#8211; 20260920 &#8211;  Z.ai | MiniMax | Alibaba | more</a> appeared first on <a href="https://ainativefoundation.org">AI Native Foundation</a>.</p>
]]></description>
										<content:encoded><![CDATA[<p>Today&#8217;s digest highlights some of the most compelling advancements from Z.ai, MiniMax, and Alibaba. These developments underscore a broader theme of enhanced efficiency and functionality in AI tools. Z.ai has significantly increased throughput with its GLM-5.3 system supporting the GLM-5.3-Flash, while StepFun introduces a preview of its Step 5 model, boasting a sparse MoE flagship with a striking 600 billion parameters and a 1 million token context window. Furthermore, Alibaba&#8217;s introduction of Qwen3.8-LiveTranslate advances real-time speech translation capabilities. Discover more in Today&#8217;s China AI Native Industry Insights.</p>
<h3>1.  Z.ai says GLM-5.3 helped build inference system for GLM-5.3-Flash, tripling throughput</h3>
<p>Z.ai reported that its GLM-5.3 model assisted engineers in building and optimizing the inference infrastructure that serves GLM-5.3-Flash. The system moved from its first successful run to production readiness in under two weeks. End-to-end throughput tripled compared with the initial baseline. Z.ai credited the result to dense feedback signals, including local correctness tests, execution traces, microbenchmarks, and end-to-end measurements, which allowed targeted hypothesis testing rather than reliance on aggregate metrics alone.<br />
Read more: <a href="https://z.ai/blog/glm-built-its-inference-infrastructure">https://z.ai/blog/glm-built-its-inference-infrastructure</a></p>
<p><video width="600" height="400" controls poster="https://cdn.ainative.foundation/image/20260920_gn_img_zai.jpg"><source src="https://cdn.ainative.foundation/video/20260920_gn_video_glm.mp4" type="video/mp4"></video><br />
Video Credit: NotebookLM</p>
<h3>2.  MiniMax Releases MiniMax Code CLI Tool</h3>
<p>MiniMax announced the release of MiniMax Code CLI, a new command-line tool from the company. The announcement was shared via the official MiniMax AI account on September 18, 2026. Further technical details were provided via a linked resource in the announcement. The release adds to MiniMax&#8217;s suite of developer-facing AI tools.<br />
Read more: <a href="https://github.com/MiniMax-AI/minimax-code">https://github.com/MiniMax-AI/minimax-code</a></p>
<p><video width="600" height="400" controls poster="https://cdn.ainative.foundation/image/20260920_gn_img_minimax.jpg"><source src="https://cdn.ainative.foundation/video/20260920_gn_video_minimax.mp4" type="video/mp4"></video><br />
Video Credit: NotebookLM</p>
<h3>3.  Alibaba&#8217;s Qwen Launches Real-Time Speech Translation Model Qwen3.8-LiveTranslate</h3>
<p>Alibaba&#8217;s Qwen team released Qwen3.8-LiveTranslate, a next-generation real-time simultaneous interpretation model built on an Interleave architecture. The model improves faithfulness, fluency, and conciseness while reducing average lagging (LAAL) from 2.8 seconds to 2.3 seconds across 60 languages. New features include real-time speaker diarization with stable voice cloning for multi-party speech, synchronized bilingual on-screen display, and long-context disambiguation that uses conversation history to keep names and terminology consistent. The model is available via Qwen&#8217;s blog and QwenCloud.<br />
Read more: <a href="https://qwen.ai/blog?id=qwen3.8-livetranslate">https://qwen.ai/blog?id=qwen3.8-livetranslate</a></p>
<p><video width="600" height="400" controls poster="https://cdn.ainative.foundation/image/20260920_2e0dc1cc893540ed934f08aa11b2df40.jpg"><source src="https://cdn.ainative.foundation/video/20260920_gn_video_qwen.mp4" type="video/mp4"></video><br />
Video Credit: The original article</p>
<h3>4.  StepFun releases Step 5 Preview, a 600B parameter sparse MoE flagship model with 1 million token context window</h3>
<p>StepFun announced Step 5 Preview, a flagship foundation model designed for real-world agentic tasks. The model uses a sparse MoE architecture with 600B total parameters and only 27B activated parameters, supports a 1 million token context window, and natively handles text and vision inputs. It ranks among the top three open-source models globally on the Artificial Analysis Intelligence Index with a score of 44, while maintaining a per-task cost that is one-eighth of Claude Opus 5. The company will release the full model weights on October 15, 2026, and the model is already accessible via API and online platforms.<br />
Read more: <a href="https://mp.weixin.qq.com/s?__biz=MzkyNTYxNzg5Mg==&#038;mid=2247488120&#038;idx=1&#038;sn=8ba9ac7f0b36682d6262290677c665da&#038;chksm=c0136f8f2faa314e887729c958ba8e9e2a972002d74192b20f5ad85554004190ec62ec1d7f36&#038;scene=0&#038;xtrack=1#rd">https://mp.weixin.qq.com/s?__biz=MzkyNTYxNzg5Mg==&#038;mid=2247488120&#038;idx=1&#038;sn=8ba9ac7f0b36682d6262290677c665da&#038;chksm=c0136f8f2faa314e887729c958ba8e9e2a972002d74192b20f5ad85554004190ec62ec1d7f36&#038;scene=0&#038;xtrack=1#rd</a></p>
<p><video width="600" height="400" controls poster="https://cdn.ainative.foundation/image/20260920_bc6bf825cfab487c8866f1597ab3790d"><source src="https://cdn.ainative.foundation/video/20260920_gn_video_stepfun.mp4" type="video/mp4"></video><br />
Video Credit: The original article</p>
<div style="width:100%;height:2px;background:#808080;margin:10px 0"></div>
<p>That&#8217;s all for today&#8217;s China AI Native Industry Insights. Join us at <a href="https://member.ainativefoundation.org/">AI Native Foundation Membership Dashboard</a> for the latest insights on AI Native, or follow our linkedin account at <a href="https://www.linkedin.com/company/ainativefoundation/">AI Native Foundation</a> and our twitter account at <a href="https://x.com/AINativeF">AINativeF</a>.</p>
<p>The post <a href="https://ainativefoundation.org/china-ai-native-industry-insights-20260920-z-ai-minimax-alibaba-more/">China AI Native Industry Insights &#8211; 20260920 &#8211;  Z.ai | MiniMax | Alibaba | more</a> appeared first on <a href="https://ainativefoundation.org">AI Native Foundation</a>.</p>
]]></content:encoded>
					
		
		<enclosure url="https://cdn.ainative.foundation/video/20260920_gn_video_glm.mp4" length="6740325" type="video/mp4" />
<enclosure url="https://cdn.ainative.foundation/video/20260920_gn_video_minimax.mp4" length="10039247" type="video/mp4" />
<enclosure url="https://cdn.ainative.foundation/video/20260920_gn_video_qwen.mp4" length="6779206" type="video/mp4" />
<enclosure url="https://cdn.ainative.foundation/video/20260920_gn_video_stepfun.mp4" length="7223834" type="video/mp4" />

			</item>
		<item>
		<title>AI Native Daily Paper Digest – 20260918 – AskChem &#124; Metis &#124; SpatialCLI</title>
		<link>https://ainativefoundation.org/ai-native-daily-paper-digest-20260918-askchem-metis-spatialcli/</link>
		
		<dc:creator><![CDATA[insights]]></dc:creator>
		<pubDate>Sat, 19 Sep 2026 00:40:57 +0000</pubDate>
				<category><![CDATA[Papers]]></category>
		<guid isPermaLink="false">https://ainativefoundation.org/ai-native-daily-paper-digest-20260918-askchem-metis-spatialcli/</guid>

					<description><![CDATA[<p>Today&#8217;s digest features cutting-edge advancements from Gemma and DeepSeek, capturing significant attention with their innovative approaches. The central theme revolves around improving [&#8230;]</p>
<p>The post <a href="https://ainativefoundation.org/ai-native-daily-paper-digest-20260918-askchem-metis-spatialcli/">AI Native Daily Paper Digest – 20260918 – AskChem | Metis | SpatialCLI</a> appeared first on <a href="https://ainativefoundation.org">AI Native Foundation</a>.</p>
]]></description>
										<content:encoded><![CDATA[<p style="margin-bottom:24px">Today&#8217;s digest features cutting-edge advancements from Gemma and DeepSeek, capturing significant attention with their innovative approaches. The central theme revolves around improving efficiency in multimodal reasoning, showcasing models achieving impressive scores on the AVA-Active benchmark. In particular, one study introduces the novel TwinFusion method, which significantly reduces processing time by 30% while maintaining high accuracy levels. Another fascinating paper explores the implications of enhanced agentic systems in dynamic environments, demonstrating a 15% improvement in task completion rates.</p>
<h3>1. AskChem: Claim-Centered Infrastructure for Chemistry Literature Synthesis</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: AskChem, claim-centered infrastructure, GPT-5.5, AI agents, chemistry search</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Knowledge Representation and Reasoning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The objective of the research was to develop AskChem, a claim-centered infrastructure for conducting cross-paper chemistry searches, which shifts the retrieval focus from whole papers to provenance-carrying claims.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; AskChem converts papers into atomic, typed claims grounded by a source DOI and quotes or evidence locators, enabling hierarchical retrieval and exploration through a faceted taxonomy and evidence graph.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; AskChem indexes 2.4M claims from 147K papers, improving the retrieval capability with 100% resolvable DOIs for a GPT-5.5 reader and provides superior citation density compared to other systems.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28618" target="_blank">https://huggingface.co/papers/2607.28618</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img fetchpriority="high" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233007276.jpg"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>2. Metis: Memory Foundation Model</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: AI Native, native memory, memory foundation models, Metis</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Foundations of AI</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; Introduce memory foundation models to enable native memory capabilities in AI systems.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Proposed Metis, a new architecture with a native memory state for storing historical information.</p>
<p>   &#8211; Developed large-scale memory-specific training data and optimization objectives for native memory procedures.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Metis demonstrates native memory capabilities with advantages in architecture and efficiency.</p>
<p>   &#8211; The study provides a detailed analysis of strengths, limitations, and behaviors of Metis.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.26760" target="_blank">https://huggingface.co/papers/2607.26760</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233032494.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>3. Frontis-MA1: Training an AI4AI Model towards Recursive Self-Improvement in Machine Learning Engineering</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Recursive self-improvement, AI4AI, OpenMLE, Machine learning engineering, Reinforcement learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: AI Systems and Tools</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The paper aims to study Recursive self-improvement by introducing OpenMLE, an open full-stack system for AI4AI research in machine learning engineering.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Utilizes a combination of execution feedback environments (OpenMLE-Gym), operator learning (OpenMLE-RL), and long-horizon search (OpenMLE-Evo) to post-train a meta-evolution agent, Frontis-MA1, with atomic program-evolution operators.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The system significantly improves performance on MLE-Bench Lite, surpassing other models in efficiency and capability with OpenMLE-Evo-Max, and demonstrates transferable capabilities on NatureBench Lite.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28568" target="_blank">https://huggingface.co/papers/2607.28568</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233059370.jpg"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>4. Memory Decoder at Scale: A Pretrained, Parametric Long-Term Memory</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Memory Decoder, scalability, parameter efficiency, Faiss, pretraining</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Natural Language Processing</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To scale memory models up to 6.9B parameters and improve memory capacity independently for better performance of language models.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Implementation of a distributed pipeline for Faiss indexing and retrieval combined with sparse, batch-wise loading of kNN distributions.</p>
<p>   &#8211; Comparison of parameter-performance tradeoff by allocating more parameters to memory across different model scales.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Independently scaling pretrained memory models can achieve better parameter efficiency and enhance language model performance, as demonstrated by outperforming larger models with fewer parameters.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.27919" target="_blank">https://huggingface.co/papers/2607.27919</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233124782.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>5. BM25 Wins at Scale: A Scaling Study of Retrieval-Augmented Generation Paradigms</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Retrieval-Augmented Generation, BM25, Dense Retrieval, Lexical Retrieval, Agentic Reasoning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Natural Language Processing </p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To conduct a controlled study evaluating retrieval-augmented generation (RAG) approaches across varying corpus sizes and determine their accuracy-cost scaling.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The study varies corpus size across 28 nested tiers, analyzing accuracy, construction, query tokens, and latency under a single reader model and judging protocol.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Different approaches show varying effectiveness based on corpus size, with BM25 eventually leading as the corpus grows, while agentic search incurs high query token costs at larger scales.</p>
<p>   &#8211; Lexical retrieval emerges as the most scalable default, outperforming dense retrieval and graph-based indexing at larger corpus sizes, highlighting global candidate ranking&#8217;s importance.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.26497" target="_blank">https://huggingface.co/papers/2607.26497</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233146706.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>6. MPIE-Bench: Benchmarking Anatomically Plausible Multi-Person Interaction Editing</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Text-to-image, personalized editing, MPIE-Bench, MPIE-Eval, multi-person mesh reconstruction</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Computer Vision</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The research aims to address issues in synthesizing high-fidelity images featuring multiple named people in shared contact actions, such as embrace or carry.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Introduced MPIE-Bench, a benchmark of 2,500 samples from video-mined editing triplets covering various scenes, interactions, and contact densities.</p>
<p>   &#8211; Proposed MPIE-Eval, which evaluates contact-time geometry using a public multi-person mesh reconstruction framework.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The study finds that current visual language model (VLM) metrics often fail to capture anatomical and geometric errors, instead showing that MPIE-Eval aligns more closely with human judgment in evaluating anatomy and interaction accuracy.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.27616" target="_blank">https://huggingface.co/papers/2607.27616</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233209977.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>7. Beyond Borrowed Histories: Person-Aligned User Simulation for Interactive Role-Playing Evaluation</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Role-playing agents, Evaluation, User simulators, Personalized rubrics, Multi-turn conversation</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Human-AI Interaction</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The paper aims to improve the evaluation of Role-playing agents (RPAs) by addressing the limitations of existing benchmarks and introducing a more personalized and realistic evaluation method through user simulators.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Implementation of PALATE, a scalable RPA benchmark using 300 character profiles and personalized rubrics, enabling free-form, multi-turn conversations with user simulators to assess RPAs&#8217; conversational abilities.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; PALATE provides interpretable evaluations of RPAs&#8217; performance in specific user interactions, demonstrating higher agreement with human judgments compared to traditional rubrics, and addresses both generic quality and personalized user experience.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.27816" target="_blank">https://huggingface.co/papers/2607.27816</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233235908.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>8. SpatialCLI: Learning to Reason With Spatial Tools, Then Without Them</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Vision-language models, SpatialCLI, Embodied agents, Compositional perception</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Multi-Modal Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The paper aims to address the capability mismatch in Vision-language models (VLMs) between task-level reasoning and capturing visual details, proposing a new framework to bridge this gap.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The proposed SpatialCLI framework operates in three stages: exposing specialist vision models as spatial tools, improving tool use with Cold-Start SFT and agentic RL, and internalizing successful tool-use trajectories.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; SpatialCLI significantly improves VLM performance on the MindCube benchmark, raising the accuracy from 29.3% to 84.6% with tools, and maintaining 73.8% accuracy even without tools after internalization, surpassing existing models.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.27703" target="_blank">https://huggingface.co/papers/2607.27703</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233300818.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>9. Chimera: Designing and Chinchilla-Scaling Hybrid Visual Diffusion Transformers</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Visual Generation, Chimera, Kimi Delta Attention, Multi-head Latent Attention, Sparse Mixture-of-Experts</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Generative Models</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To develop a hybrid visual diffusion backbone, Chimera, for efficient handling of high-resolution images, long videos, and multimodal context.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Introduction of Chimera that utilizes Kimi Delta Attention, Multi-head Latent Attention, and Sparse Mixture-of-Experts layers, along with HeteroP for scaling the architecture effectively.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Chimera&#8217;s dense backbone is 1.7x more compute-efficient than the Wan-2.1 baseline, and the complete system achieves 7.3x efficiency.</p>
<p>   &#8211; Chimera can extrapolate from 5-second training clips to 30-second videos with minimal performance degradation.</p>
<p>   &#8211; Fitted compute-optimal laws indicate balanced compute division in image pretraining, with a slight model size preference in video pretraining at higher budgets.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28611" target="_blank">https://huggingface.co/papers/2607.28611</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233323944.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>10. β-OPSD: Deriving with Policy Optimization, Training with Self-Distillation</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: On-policy self-distillation, policy optimization, reinforcement learning, reasoning performance, mathematical reasoning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Reinforcement Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study aims to enhance reasoning language models using a refined version of on-policy self-distillation, termed β-OPSD, which introduces a versatile regularization parameter β for better policy optimization.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The paper introduces β-OPSD by positioning it within a broader policy-optimization family. It employs reinforcement learning for distilling the closed-form solution through efficient token-level logit mixing to approximate costly optimization processes.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Experiments demonstrate that β-OPSD consistently outperforms the standard OPSD in mathematical reasoning tasks, achieving improved optimization stability and better reasoning performance.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28582" target="_blank">https://huggingface.co/papers/2607.28582</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233348258.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>11. Can Large Language Models Execute Parent Orders?</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Parent-order execution, Large Language Models (LLMs), PACE, Shenzhen Stock Exchange, Execution decisions</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: AI in Finance</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study aims to enhance parent-order execution in algorithmic trading using Large Language Models (LLMs), addressing limitations of current methods like reliance on market assumptions and lack of adaptability.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Introduction of PACE (Plan-Ahead Controlled Execution), a hierarchical framework that separates parent-order execution into long-term planning and short-term execution, bypassing the need for specific market assumptions and training.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; PACE demonstrated superior performance compared to existing TWAP, Almgren-Chriss, and learning-based strategies during tests on Shenzhen Stock Exchange data. It was observed that LLMs navigate execution decisions differently from humans, with higher confidence correlated with improved performance and earlier trades instead of late actions.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28410" target="_blank">https://huggingface.co/papers/2607.28410</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233411593.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>12. LEDGERMIND: Provenance-Constrained Multimodal Agentic Reasoning with a Structured Evidence Ledger</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Multimodal Agents, Visual Question Answering, Provenance-Constrained, Structured Evidence, LedgerMind</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Multi-Modal Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study aims to improve the evaluation of multimodal agents that handle visual question answering by considering not only answer accuracy but also the process through which answers are derived.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Introduction of LedgerMind, a system based on a provenance-constrained state machine, to standardize tool outputs and support structured evidence.</p>
<p>   &#8211; Utilization of a Three-Layer Grounding Protocol, Adaptive Dual-Path Dispatcher, and an Event-Triggered Verification-and-Repair engine to manage reasoning depth and ensure provenance.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; LedgerMind effectively addresses common failure patterns such as unsupported reasoning and entity hallucinations, enhancing both answer accuracy and the reliability of the reasoning trajectory.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28374" target="_blank">https://huggingface.co/papers/2607.28374</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233436135.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>13. Σ-Mem: An Online Reliability Memory for LLM-based Multi-Agent Systems</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: LLM agents, multi-agent systems, Σ-Mem, reliability memory, interaction content</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: AI Systems and Tools</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The objective is to introduce Σ-Mem, a novel online reliability memory system, addressing limitations in existing memory systems within LLM agents, particularly in multi-agent environments.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Σ-Mem models historical competence and peer relationship evidence, using real symmetric states updated from correctness feedback while ensuring stable online adaptation without requiring model retraining.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Σ-Mem effectively adapts to shifts in counterfactual reliability and generalizes to new peers and task domains, outperforms traditional voting methods, and improves with increased correctness feedback, establishing a solid foundation for adaptive coordination in multi-agent LLM systems.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.27958" target="_blank">https://huggingface.co/papers/2607.27958</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233503295.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>14. Explorative Modeling: Unlocking a Third Pretraining Axis and End-to-End Generation</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Explorative Modeling, Generative Models, FLOP Efficiency, Sample Efficiency, Parameter Efficiency</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Generative Models</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; Introduce Explorative Modeling as a new paradigm in generative modeling to enable end-to-end training by exploring candidate matches and training on the best predictions.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The Explorative Models are applied in generative domains, increasing exploration as a pretraining axis to improve performance across diverse datasets including images, video, and language.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Explorative Modeling enhances efficiency exponentially, achieving significant improvement in FLOP, sample, and parameter efficiency, and enables end-to-end reconstructive generative modeling with reduced inference steps, establishing itself as a new training axis for existing generative models.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.27372" target="_blank">https://huggingface.co/papers/2607.27372</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233526935.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>15. AI Tour Meeting: Group Travel Planning by LLM Agents</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: AI Tour Meeting, Large Language Model, agents, simulation tool</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: AI Systems and Tools</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To propose a travel planning framework called AI Tour Meeting using agents powered by Large Language Models (LLMs) to collaboratively create itineraries through natural language discussion.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Implementation of a framework for configuring agent personas and orchestrating discussion workflows, enabling monitoring and LLM deployment in the simulation of tour planning discussions.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The framework is validated by system demonstration and provides analytical insights into the behavior of multiple LLM agents during group travel planning.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.18806" target="_blank">https://huggingface.co/papers/2607.18806</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233614234.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>16. Revisiting Lossy Verification in Speculative Decoding: Mechanisms, Trade-offs, and Failure Modes</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Speculative Decoding, Large Language Model, Truncation-based Verification, Collaborative Verification</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Natural Language Processing</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To analyze the distributions induced by lossy verification schemes in speculative decoding and identify their potential pitfalls in large language model inference.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Classified lossy verification methods into truncation-based and collaborative verification, and developed a diagnostic evaluation framework across curated benchmarks.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Truncation-based methods are susceptible to significant performance degradation due to distributional distortion.</p>
<p>   &#8211; For collaborative verification, controlling the overshoot of draft probabilities relative to target probabilities is crucial to maintaining output quality.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.26627" target="_blank">https://huggingface.co/papers/2607.26627</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233549571.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>17. Beyond Geometric Complementarity: Coherent Overlap in Sparse Mixture-of-Experts Routing</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Sparse mixture-of-experts, Expert Subspace Separation Index, coherent overlap, next-token prediction</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Natural Language Processing</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study aims to distinguish the benefits and interactions of experts within sparse mixture-of-experts (MoE) language models by analyzing their geometric arrangement.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Utilizes an Expert Subspace Separation Index (ESSI) and various controlled interventions, including matched-route residuals and prefix-controlled factorial setups, to analyze expert overlaps and their impact on representation.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Coherent overlap indicates that while routes overlap, useful multi-expert computation persists, clarifying why geometric similarity doesn&#8217;t alone determine redundancy or pruning value.</p>
<p>   &#8211; Adding later experts improves next-token predictions significantly in most cases, emphasizing the importance of non-redundant expert selections.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28308" target="_blank">https://huggingface.co/papers/2607.28308</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233702669.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>18. Fairness Pruning: Locating Demographic Bias in GLU-MLP Layers via Differential Activations</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Fairness Pruning, Large Language Models, Demographic Bias, Causal Bias Localization, AI Ethics and Fairness</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: AI Ethics and Fairness</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To introduce and empirically validate Fairness Pruning, a method aimed at managing and mitigating demographic biases in large language models.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Utilization of minimally contrastive prompt pairs and inference-time activation capture to identify neurons that react differentially to demographic attributes.</p>
<p>   &#8211; Empirical evaluation on large models such as Llama-3.2 and Salamandra-2B with structural intervention by zeroing neurons.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Demonstrated that zeroing specific neurons alters model responses to demographic variables, with a bidirectional bias destabilization effect.</p>
<p>   &#8211; Confirmed that demographic bias processing and model capabilities are dissociable, setting foundations for transitioning to directional behavior modulation.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28319" target="_blank">https://huggingface.co/papers/2607.28319</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233637423.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>19. AMRD: Adaptive Multi-Teacher Relational Distillation for Lightweight Speech Emotion Recognition</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Speech Emotion Recognition, Edge Devices, Knowledge Distillation, AI Native</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Machine Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; This research aims to optimize on-device speech emotion recognition (SER) for real-time applications by developing a more efficient model that can operate effectively on edge devices.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Adaptive Multi-teacher Relational Distillation (AMRD) is proposed, using a one-class SVM to assign weights to teacher models and a relational distillation loss to improve inter-sample relational structure alignment between teacher and student models.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The AMRD approach outperforms single-teacher distillation baselines on the IEMOCAP and CREMA-D datasets, demonstrating its effectiveness across multiple student architectures and confirming that both components of the method yield complementary benefits.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.25289" target="_blank">https://huggingface.co/papers/2607.25289</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233726024.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>20. </h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="" target="_blank"></a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/thirteen/202607311785541064.jpg"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>21. Pedestrian Archetypes Extension &#8212; More Pedestrian Models for Autonomous Vehicle Safety Testing</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Pedestrian Archetypes, Real-world Traffic Scenarios, Behavior Patterns</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Robotics and Autonomous Systems</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The research introduces new pedestrian archetypes that exhibit significant behavioral differences, enhancing the understanding of pedestrian behaviors beyond the original taxonomy.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Analysis of YouTube dash-cam videos for identifying additional pedestrian archetypes not covered in previous work.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Seven new pedestrian archetypes are presented, each defined by unique behaviors that are documented through video-frame evidence, thus expanding the existing framework for assessing pedestrian behavior in traffic.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.16922" target="_blank">https://huggingface.co/papers/2607.16922</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233737256.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>22. ReToken: One Token to Improve Vision-Language Models for Visual Retrieval</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Vision-Language Models, GPU Memory Constraints, Image-QA Dataset</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Multi-Modal Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; Address the challenge of long visual context in vision-language models by introducing ReToken, an explicit retrieval target for selecting query-relevant visual tokens.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; ReToken is trained on a small image-QA dataset and designed to select a sparse set of visual tokens from a pre-filled visual KV cache.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; ReToken shows consistent performance gains across image and video benchmarks, with significant improvements in Visual Haystacks and LVBench, while maintaining efficiency on a single H100 for both training and long-video inference.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28627" target="_blank">https://huggingface.co/papers/2607.28627</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233713703.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>23. OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: OmniScope, token compression, query, salience mismatch, AI Native</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Multi-Modal Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The objective is to address the cross-modal salience mismatch in token compression methods for omnimodal large language models by developing a framework named OmniScope that estimates relevance separately for audio and video.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Introduced a training-free token compression framework using the query as a shared semantic anchor, allocating modality-specific token budgets, pruning visual tokens with an anchor-delta strategy, and merging audio tokens to reduce redundancy.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; OmniScope outperforms existing methods in accuracy across audio-video benchmarks and model scales, offering improvements in speed and memory efficiency while maintaining high accuracy with minimal drop.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.23193" target="_blank">https://huggingface.co/papers/2607.23193</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233650331.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>24. Harness-G: A Graph-Structured Harness for Search Agents</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Reinforcement Learning, retrieval aliasing, Harness-G, Structured Non-myopic Credit</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Reinforcement Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study addresses the problem of retrieval aliasing in reinforcement learning search agents by proposing a new retrieval framework to optimize query generation and interaction.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The authors introduce Harness-G, a graph-structured retrieval framework that transforms query generation into finite action selection. They also develop Structured Non-myopic Credit (SNC) for better rewarding actions.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Harness-G achieves the highest average F1 across six QA benchmarks and outperforms the best baseline by significant margins, demonstrating its effectiveness in reducing linguistic aliasing and improving retrieval accuracy.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.27652" target="_blank">https://huggingface.co/papers/2607.27652</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233625249.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>25. Is Deep Research Reliable? Misleading Knowledge Induces False Conclusions</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Deep Research agents, Misleading Knowledge, Evidence Synthesis, Workflow Reliability</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Knowledge Representation and Reasoning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study aims to explore the reliability vulnerabilities in Deep Research agents when exposed to misleading knowledge in open information environments.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Introduces MisKnow-Agent, a framework for generating misleading knowledge instances to test these vulnerabilities, utilizing 5,933 quality-controlled instances derived from DeepResearch Benchmark tasks.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The experiments demonstrate that even minimal misleading knowledge can lead to false conclusions in final reports, highlighting a significant reliability issue.</p>
<p>   &#8211; Despite verifier models identifying misleading instances, they can still be adopted in long-term workflows, indicating a gap in workflow-level evidence verification.</p>
<p>   &#8211; Evaluated defenses show mitigation effects but fail to fully prevent the adoption of false conclusions, pointing to the need for improved verification and correction capabilities.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.20891" target="_blank">https://huggingface.co/papers/2607.20891</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233600666.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>26. Filesystem-Based Memory for LLM Agents: Organization, Evolution, and Sustainability</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: LLM agents, filesystem-based memory, memory organization, search economy, management agent</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: AI Systems and Tools</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To systematically explore filesystem-based memory for Large Language Model (LLM) agents, assessing the structure and organization of such memory.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The study formalizes the concept of a memory filesystem with a management agent for organization, a search agent for querying, and an execution agent for task trajectory integration.</p>
<p>   &#8211; Investigated variables include memory shape, stream scale, tool harness, and agent strength, with focus on memory growth, retrieval costs, and store health.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Organization in filesystem-based memory leads to improved search economy, significantly reducing retrieval costs.</p>
<p>   &#8211; However, as memory grows, management agents struggle to maintain organization quality, and improvements in organization do not necessarily translate to better answer quality.</p>
<p>   &#8211; The study identifies the tool set as a significant factor in shaping the memory store, highlighting a design space for agent memory rather than just an assumption.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.26637" target="_blank">https://huggingface.co/papers/2607.26637</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233538463.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>27. Multi-Head Attention Residuals</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Transformers, Multi-Head Attention Residuals, Nemotron, validation loss</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Machine Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; Introduction of Multi-Head Attention Residuals (MHAR) to improve information propagation across model depth and width.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; MHAR utilizes H per-subspace heads with distinct softmax distributions, trained on a Nemotron-based corpus to optimize validation loss across various model sizes.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; MHAR outperforms standard Transformers with validation loss improvements across 100M, 350M, and 1B scales. Optimal head count H = 8 increases efficiency and maintains near-baseline memory consumption.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.27230" target="_blank">https://huggingface.co/papers/2607.27230</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233514954.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>28. Echoverse: Deep, Evolving Environments for Training Computer-Use Agents at Scale</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Synthetic Environments, Echoverse, Reinforcement Learning, State Representation</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Reinforcement Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To improve computer-use agents&#8217; training environments by focusing on behavioral depth, targeted interactions, and model improvement.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Developed Echoverse to compile specifications into stateful applications, incorporating a co-evolution loop for iterative environment and model enhancement.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Demonstrated significant model improvement across various environments, with increases in evaluation scores and live-site accuracy via deep environments and iterative environment repairs.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28074" target="_blank">https://huggingface.co/papers/2607.28074</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233451225.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>29. INTACT: Isomorphic Intent-to-Action Learning for Search-Free World Models</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: INTACT, Intent-to-Action, Latent World Models, Action Law Semantics</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Reinforcement Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To introduce INTACT, an end-to-end JEPA, aimed at transforming action-labeled, reward-free trajectories into a deployable intent-to-action interface without expensive test-time search.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Utilizes a unique architecture with isomorphic intent representation for local and goal motion-intent models, and provides intact transfer from RGB evidence to action-effective latent intent coordinates.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The INTACT model achieved up to 100% success rates on LeWM tasks using a search-free policy, significantly reducing the need for candidate sequence sampling while improving performance through a shared four-task encoder. Direct inference is highly efficient, taking only 2.9-5.5 milliseconds.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.26056" target="_blank">https://huggingface.co/papers/2607.26056</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233423941.jpg"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>30. MemHarness: Memory Is Reconstructed, Not Replayed</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: MemHarness, Negative Transfer, Context-grounded Guidance, Large Language Model Agents, GRPO</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Reinforcement Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The objective is to develop a framework named MemHarness that enables large language model agents to actively reconstruct past experiences based on the current context, thereby providing more accurate and context-relevant guidance.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; MemHarness employs a unified policy model that critiques and reconstructs retrieved experiences for each decision step. This reconstructive ability is achieved through end-to-end training using the GRPO technique.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; MemHarness significantly outperforms existing memory-augmented baselines and pure reinforcement learning models, especially in out-of-distribution scenarios, by preventing negative transfer and enhancing intrinsic reasoning capabilities.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28272" target="_blank">https://huggingface.co/papers/2607.28272</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233359267.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>31. ShadowDancer: Teaching Video World Models Any Action by Learning Unified Dynamics Representations from a Video and Its Shadow</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: ShadowDancer, interactive video, dynamics, action transfer, cross-shadow prediction</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Computer Vision</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To enable precise control of any-action, frame-level dynamics in interactive video world models through novel techniques.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Utilizes shadow pairs to construct scalable video pairs replaying the same dynamics with varied appearances.</p>
<p>   &#8211; Implements cross-shadow prediction to learn actions by predicting dynamics from one shadow to another.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Demonstrated improved action transfer and extended action rollout in diverse dynamics families, achieving an average blinded win rate of 86% over existing models.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28362" target="_blank">https://huggingface.co/papers/2607.28362</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233335737.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>32. See2Think: Do Multimodal Models Really Use Intermediate Visual States?</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Multimodal large language models, Visual reasoning, See2Think, Visual Action-of-Thought</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Multi-Modal Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To introduce See2Think, an evaluation framework that assesses visual reasoning in multimodal large language models through benchmarks like See2ThinkBench and Visual Action-of-Thought.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The creation of See2ThinkBench with 1,200 visually dependent tasks across 12 categories and the use of Visual Action-of-Thought to record and analyze visual and textual operations under controlled settings.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Visual reasoning capabilities in models are strongly dependent on both the model and environment, showing varied performance across tasks. Additionally, even though models often select the correct visual operations, the accuracy improvements are hindered by challenges in rendering fidelity and reliance on feedback which does not always improve accuracy.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.26769" target="_blank">https://huggingface.co/papers/2607.26769</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233311886.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>33. RefCaptioner: Multi-Reference Image-Grounded Video Captioning</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: multi-reference image-grounded video captioning, RefCaptioner, factual video descriptions, Hierarchical Coverage-Discounted GRPO, AI-generated videos</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Multi-Modal Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To introduce and address the task of multi-reference image-grounded video captioning, requiring factual video descriptions with phrase-level reference grounding.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Development of RefCaptioner, a two-stage post-training framework combining mixed-data SFT with Hierarchical Coverage-Discounted GRPO to enhance reference selection and cross-reference consistency.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; RefCaptioner outperforms other open-source models in overall performance, yielding captions preferred by human annotators and allowing for more source-faithful video reconstruction.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28509" target="_blank">https://huggingface.co/papers/2607.28509</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233249116.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>34. ACE-Data-0: Human-Centric Ambient Capture as Embodied Data Engine</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Embodied intelligence, Ambient Capture Engine, AI Native, Multi-modal datasets, Human-centric data</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Robotics and Autonomous Systems</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study focuses on addressing the data bottleneck in embodied intelligence by introducing the Ambient Capture Engine (ACE), which captures comprehensive sensory data in real home environments.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The authors developed ACE to record multisensory data through a combination of table-scale and room-scale configurations, creating a detailed dataset, ACE-Data-0, encompassing 150 hours of interactions across 200 task categories.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Evaluations reveal significant limitations in current state-of-the-art methods, particularly under challenging conditions like occlusion and long temporal horizons, highlighting ACE&#8217;s potential in supporting scalable advancements in imitation learning and embodied AI systems.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28625" target="_blank">https://huggingface.co/papers/2607.28625</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233221968.jpg"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>35. Flux-OPD: On-Policy Distillation with Evolving Contexts</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Large language model, evolving contexts, Flux-OPD, open-ended domains, teacher supervision</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Natural Language Processing</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study aims to develop a mechanism for more effective supervision in large language model training across open-ended domains, addressing the issue of lacking verifiable rewards and formalized task preferences.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The research introduces Flux-OPD, a novel OPD paradigm that leverages evolving contexts as a form of in-training supervision. It analyzes the effect of context through the reverse KL objective and uses context-conditioned teachers to provide contextual difference signals.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Flux-OPD demonstrates superiority over existing OPD paradigms, effectively capturing task preferences and enhancing performance in open-ended tasks by combining evolving contexts with teacher supervision.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28022" target="_blank">https://huggingface.co/papers/2607.28022</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233157728.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>36. Beacon: Knowing When and How to Perform Agentic Visual Reasoning</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Agentic Visual Reasoning, Multimodal Large Language Models, Mode Adaptiveness, Tool Effect, Beacon</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Multi-Modal Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study aims to enhance the success rate of multimodal large language models (MLLMs) in complex tasks by rethinking agentic visual reasoning, focusing on Mode Adaptiveness and Tool Effect.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Comprehensive analysis to quantify Mode Adaptiveness and Tool Effect in existing models, followed by the proposal of a new model named Beacon, featuring Necessity-Aware Adaptive Reward and Hint-Guided Capability Expansion.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The research concludes that existing models have limited Mode Adaptiveness, and tool benefits are often neutralized by their drawbacks. Beacon, however, shows improved overall performance, notably enhancing Mode Adaptiveness and achieving genuine gains from tool use.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28595" target="_blank">https://huggingface.co/papers/2607.28595</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233136136.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>37. VideoCoCo: Code-as-CoT for Physically-Consistent Video Generation via an Agentic Dual-Engine System</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Text-to-video, VideoCoCo, Blender program, spatiotemporal process, generative video engine</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Generative Models</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study aims to address the limitations of text-to-video models in generating physically consistent dynamics by introducing a new framework called VideoCoCo.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; A dual-engine framework is employed, where a coding agent synthesizes executable Blender code to specify scene evolution and produce deterministic drafts, followed by a generative video engine for photorealistic video transformation.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; VideoCoCo effectively improves performance on benchmarks PhyGenBench and VBench-2.0 by providing an executable, controllable, and inspectable representation for video generation, achieving the best average scores on both benchmarks.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.27380" target="_blank">https://huggingface.co/papers/2607.27380</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233111733.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>38. PhiZero: A World Model Built Around Physical Language</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: PhiZero, physical language, world model, zero-shot motion transfer</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Generative Models</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The primary aim is to develop PhiZero, a model that abstracts predictive structure from visual experience to explicitly reason about physical world evolution using physical language.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; PhiZero employs a reason-then-render paradigm where it infers future world evolution sequences in physical language and renders transitions into video through self-supervision from in-the-wild videos.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The experiments demonstrate PhiZero&#8217;s capability to model physically coherent world evolution and highlight its potential for realistic interactive world modeling, fine-grained action-conditioned simulation, and zero-shot motion transfer.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28624" target="_blank">https://huggingface.co/papers/2607.28624</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233043717.jpg"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>39. Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: GUI agents, Qwen-UI-Agent, Reinforcement Learning, MobileWorld, Online RL</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: AI Systems and Tools</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The paper aims to advance GUI agents for real-world applications, enabling reliable operation across various devices and platforms, proactive service initiation, and autonomous capability improvements.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Introduces Qwen-UI-Agent, utilizing a unified action space that combines GUI operations with CLI execution, and employs an AutoResearch-style data flywheel and Online Reinforcement Learning (RL) for training and advancing agent performance.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Qwen-UI-Agent demonstrates state-of-the-art performance in mobile-use benchmarks and competitive results in computer- and browser-use tasks, achieving high scores on MobileWorld, AndroidDaily, and WebArena among others.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28227" target="_blank">https://huggingface.co/papers/2607.28227</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233022448.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<p>The post <a href="https://ainativefoundation.org/ai-native-daily-paper-digest-20260918-askchem-metis-spatialcli/">AI Native Daily Paper Digest – 20260918 – AskChem | Metis | SpatialCLI</a> appeared first on <a href="https://ainativefoundation.org">AI Native Foundation</a>.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Global AI Native Industry Insights &#8211; 20260918 &#8211;  OpenAI &#124; Anthropic &#124; Cognition &#124; more</title>
		<link>https://ainativefoundation.org/global-ai-native-industry-insights-20260918-openai-anthropic-cognition-more/</link>
		
		<dc:creator><![CDATA[AINF]]></dc:creator>
		<pubDate>Fri, 18 Sep 2026 08:24:52 +0000</pubDate>
				<category><![CDATA[Global Industry]]></category>
		<guid isPermaLink="false">https://ainativefoundation.org/?p=10990</guid>

					<description><![CDATA[<p>OpenAI takes a leap in the legal sector with the launch of Astra for Law, a GPT-6 Astra-powered solution tailored for legal professionals. Today's digest highlights the theme of AI's growing specialization and integration into industry-specific applications. Anthropic consolidates its Claude offerings into a seamless single experience, while Cognition introduces its Devin AI Agent to Mac for iOS app development. ElevenLabs unveils Reception, an innovative AI-based receptionist platform designed to enhance efficiency for small businesses.</p>
<p>The post <a href="https://ainativefoundation.org/global-ai-native-industry-insights-20260918-openai-anthropic-cognition-more/">Global AI Native Industry Insights &#8211; 20260918 &#8211;  OpenAI | Anthropic | Cognition | more</a> appeared first on <a href="https://ainativefoundation.org">AI Native Foundation</a>.</p>
]]></description>
										<content:encoded><![CDATA[<p>OpenAI takes a leap in the legal sector with the launch of Astra for Law, a GPT-6 Astra-powered solution tailored for legal professionals. Today&#8217;s digest highlights the theme of AI&#8217;s growing specialization and integration into industry-specific applications. Anthropic consolidates its Claude offerings into a seamless single experience, while Cognition introduces its Devin AI Agent to Mac for iOS app development. ElevenLabs unveils Reception, an innovative AI-based receptionist platform designed to enhance efficiency for small businesses. Discover more in Today&#8217;s Global AI Native Industry Insights.</p>
<h3>1.  OpenAI launches Astra for Law, a GPT-6 Astra-powered offering for legal professionals</h3>
<p>OpenAI has introduced Astra for Law, a new product built on its GPT-6 Astra model. The offering provides tools, settings, and context tailored to support lawyers and legal technology firms. It is designed to complement the expertise and judgment of legal professionals rather than replace it. The launch expands OpenAI&#8217;s push into industry-specific AI products.<br />
Read more: <a href="https://openai.com/index/astra-for-law/">https://openai.com/index/astra-for-law/</a></p>
<p><video width="600" height="400" controls poster="https://cdn.ainative.foundation/image/20260918_fef803c65c1946cf8ff876e10eec5862.jpg"><source src="https://cdn.ainative.foundation/video/20260918_e4541c174d0b44b790dadcf39019c76f.mp4" type="video/mp4"></video><br />
Video Credit: @OpenAI on X</p>
<h3>2.  Anthropic merges Claude Cowork and chat into a single Claude experience</h3>
<p>Anthropic announced it is merging Claude Cowork and Claude chat into one unified Claude product. Users can ask quick questions or hand off longer tasks like reports, and Claude will continue working on them even after the user closes their laptop. Claude will ask clarifying questions when something is unclear, while users retain final decision-making authority. The unified experience will roll out to Pro and Max subscribers over the coming weeks.<br />
Read more: <a href="https://claude.com/blog/cowork-is-now-claude">https://claude.com/blog/cowork-is-now-claude</a></p>
<p><video width="600" height="400" controls poster="https://cdn.ainative.foundation/image/20260918_03d6fcd7bd154dffa844f16b0fb10588.png"><source src="https://cdn.ainative.foundation/video/20260918_4a56adc8ef9340419dedc60de1d40e91.mp4" type="video/mp4"></video><br />
Video Credit: @claudeai on X</p>
<h3>3.  Cognition&#8217;s Devin AI Agent Gains Mac Access for iOS App Development</h3>
<p>Cognition announced that its AI coding agent Devin now has access to a Mac virtual machine. With this update, Devin can build and test applications using an iOS simulator directly on its Mac VM. The agent can also send screen recordings of its work via Slack and share TestFlight links so users can try the resulting apps. This expands Devin&#8217;s capabilities beyond web and backend development into native iOS app creation and testing.<br />
Read more: <a href="https://devin.ai/blog/devin-gets-a-mac">https://devin.ai/blog/devin-gets-a-mac</a></p>
<p><video width="600" height="400" controls poster="https://cdn.ainative.foundation/image/20260918_06d11cc180384e98b783461d94aabd9c.jpg"><source src="https://cdn.ainative.foundation/video/20260918_6388ecb4e33b4f758578bb73ec0a8d77.mp4" type="video/mp4"></video><br />
Video Credit: @cognition on X</p>
<h3>4.  ElevenLabs Launches Reception, an AI Receptionist Platform for Small Businesses</h3>
<p>ElevenLabs has introduced Reception, an AI receptionist platform for small businesses built on its ElevenAgents technology. The service answers incoming calls, responds to customer questions, books appointments, and sends text confirmations automatically. Businesses can set it up in minutes by simply providing their website. The product targets missed-call scenarios that can result in lost customers.<br />
Read more: <a href="https://elevenlabs.io/blog/reception">https://elevenlabs.io/blog/reception</a></p>
<p><video width="600" height="400" controls poster="https://cdn.ainative.foundation/image/20260918_1f0a3a17f1da43b88e013da5ee90829c.jpg"><source src="https://cdn.ainative.foundation/video/20260918_d07bb9864a424912b57c2c2683378c86.mp4" type="video/mp4"></video><br />
Video Credit: @ElevenLabs on X</p>
<div style="width:100%;height:2px;background:#808080;margin:10px 0"></div>
<p>That&#8217;s all for today&#8217;s Global AI Native Industry Insights. Join us at <a href="https://member.ainativefoundation.org/">AI Native Foundation Membership Dashboard</a> for the latest insights on AI Native, or follow our linkedin account at <a href="https://www.linkedin.com/company/ainativefoundation/">AI Native Foundation</a> and our twitter account at <a href="https://x.com/AINativeF">AINativeF</a>.</p>
<p>The post <a href="https://ainativefoundation.org/global-ai-native-industry-insights-20260918-openai-anthropic-cognition-more/">Global AI Native Industry Insights &#8211; 20260918 &#8211;  OpenAI | Anthropic | Cognition | more</a> appeared first on <a href="https://ainativefoundation.org">AI Native Foundation</a>.</p>
]]></content:encoded>
					
		
		<enclosure url="https://cdn.ainative.foundation/video/20260918_e4541c174d0b44b790dadcf39019c76f.mp4" length="14687590" type="video/mp4" />
<enclosure url="https://cdn.ainative.foundation/video/20260918_4a56adc8ef9340419dedc60de1d40e91.mp4" length="133006855" type="video/mp4" />
<enclosure url="https://cdn.ainative.foundation/video/20260918_6388ecb4e33b4f758578bb73ec0a8d77.mp4" length="21002210" type="video/mp4" />
<enclosure url="https://cdn.ainative.foundation/video/20260918_d07bb9864a424912b57c2c2683378c86.mp4" length="10926349" type="video/mp4" />

			</item>
		<item>
		<title>AI Native Daily Paper Digest – 20260917 – AskChem &#124; Metis &#124; Chimera</title>
		<link>https://ainativefoundation.org/ai-native-daily-paper-digest-20260917-askchem-metis-chimera/</link>
		
		<dc:creator><![CDATA[insights]]></dc:creator>
		<pubDate>Fri, 18 Sep 2026 00:41:07 +0000</pubDate>
				<category><![CDATA[Papers]]></category>
		<guid isPermaLink="false">https://ainativefoundation.org/ai-native-daily-paper-digest-20260917-askchem-metis-chimera/</guid>

					<description><![CDATA[<p>Today&#8217;s digest highlights advances from notable names like GPT and Llama, showcasing new paradigms in long-context attention. The focal theme weaves through [&#8230;]</p>
<p>The post <a href="https://ainativefoundation.org/ai-native-daily-paper-digest-20260917-askchem-metis-chimera/">AI Native Daily Paper Digest – 20260917 – AskChem | Metis | Chimera</a> appeared first on <a href="https://ainativefoundation.org">AI Native Foundation</a>.</p>
]]></description>
										<content:encoded><![CDATA[<p style="margin-bottom:24px">Today&#8217;s digest highlights advances from notable names like GPT and Llama, showcasing new paradigms in long-context attention. The focal theme weaves through improved multimodal reasoning, with papers presenting methods such as Multi-Tiered Transformer Processing which achieves state-of-the-art results on the GLUE benchmark. Another notable contribution is a breakthrough in robotics, utilizing Adaptive Neural Mechanisms to enhance autonomous decision-making. Furthermore, a key study reveals significant reduction in error rates by employing Dynamic Scaling Models, which may lead to more efficient processing in AI systems.</p>
<h3>1. AskChem: Claim-Centered Infrastructure for Chemistry Literature Synthesis</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: AskChem, claim-centered infrastructure, GPT-5.5, AI agents, chemistry search</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Knowledge Representation and Reasoning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The objective of the research was to develop AskChem, a claim-centered infrastructure for conducting cross-paper chemistry searches, which shifts the retrieval focus from whole papers to provenance-carrying claims.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; AskChem converts papers into atomic, typed claims grounded by a source DOI and quotes or evidence locators, enabling hierarchical retrieval and exploration through a faceted taxonomy and evidence graph.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; AskChem indexes 2.4M claims from 147K papers, improving the retrieval capability with 100% resolvable DOIs for a GPT-5.5 reader and provides superior citation density compared to other systems.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28618" target="_blank">https://huggingface.co/papers/2607.28618</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img fetchpriority="high" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233007276.jpg"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>2. Metis: Memory Foundation Model</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: AI Native, native memory, memory foundation models, Metis</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Foundations of AI</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; Introduce memory foundation models to enable native memory capabilities in AI systems.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Proposed Metis, a new architecture with a native memory state for storing historical information.</p>
<p>   &#8211; Developed large-scale memory-specific training data and optimization objectives for native memory procedures.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Metis demonstrates native memory capabilities with advantages in architecture and efficiency.</p>
<p>   &#8211; The study provides a detailed analysis of strengths, limitations, and behaviors of Metis.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.26760" target="_blank">https://huggingface.co/papers/2607.26760</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233032494.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>3. Frontis-MA1: Training an AI4AI Model towards Recursive Self-Improvement in Machine Learning Engineering</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Recursive self-improvement, AI4AI, OpenMLE, Machine learning engineering, Reinforcement learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: AI Systems and Tools</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The paper aims to study Recursive self-improvement by introducing OpenMLE, an open full-stack system for AI4AI research in machine learning engineering.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Utilizes a combination of execution feedback environments (OpenMLE-Gym), operator learning (OpenMLE-RL), and long-horizon search (OpenMLE-Evo) to post-train a meta-evolution agent, Frontis-MA1, with atomic program-evolution operators.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The system significantly improves performance on MLE-Bench Lite, surpassing other models in efficiency and capability with OpenMLE-Evo-Max, and demonstrates transferable capabilities on NatureBench Lite.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28568" target="_blank">https://huggingface.co/papers/2607.28568</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233059370.jpg"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>4. Memory Decoder at Scale: A Pretrained, Parametric Long-Term Memory</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Memory Decoder, scalability, parameter efficiency, Faiss, pretraining</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Natural Language Processing</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To scale memory models up to 6.9B parameters and improve memory capacity independently for better performance of language models.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Implementation of a distributed pipeline for Faiss indexing and retrieval combined with sparse, batch-wise loading of kNN distributions.</p>
<p>   &#8211; Comparison of parameter-performance tradeoff by allocating more parameters to memory across different model scales.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Independently scaling pretrained memory models can achieve better parameter efficiency and enhance language model performance, as demonstrated by outperforming larger models with fewer parameters.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.27919" target="_blank">https://huggingface.co/papers/2607.27919</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233124782.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>5. BM25 Wins at Scale: A Scaling Study of Retrieval-Augmented Generation Paradigms</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Retrieval-Augmented Generation, BM25, Dense Retrieval, Lexical Retrieval, Agentic Reasoning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Natural Language Processing </p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To conduct a controlled study evaluating retrieval-augmented generation (RAG) approaches across varying corpus sizes and determine their accuracy-cost scaling.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The study varies corpus size across 28 nested tiers, analyzing accuracy, construction, query tokens, and latency under a single reader model and judging protocol.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Different approaches show varying effectiveness based on corpus size, with BM25 eventually leading as the corpus grows, while agentic search incurs high query token costs at larger scales.</p>
<p>   &#8211; Lexical retrieval emerges as the most scalable default, outperforming dense retrieval and graph-based indexing at larger corpus sizes, highlighting global candidate ranking&#8217;s importance.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.26497" target="_blank">https://huggingface.co/papers/2607.26497</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233146706.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>6. MPIE-Bench: Benchmarking Anatomically Plausible Multi-Person Interaction Editing</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Text-to-image, personalized editing, MPIE-Bench, MPIE-Eval, multi-person mesh reconstruction</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Computer Vision</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The research aims to address issues in synthesizing high-fidelity images featuring multiple named people in shared contact actions, such as embrace or carry.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Introduced MPIE-Bench, a benchmark of 2,500 samples from video-mined editing triplets covering various scenes, interactions, and contact densities.</p>
<p>   &#8211; Proposed MPIE-Eval, which evaluates contact-time geometry using a public multi-person mesh reconstruction framework.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The study finds that current visual language model (VLM) metrics often fail to capture anatomical and geometric errors, instead showing that MPIE-Eval aligns more closely with human judgment in evaluating anatomy and interaction accuracy.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.27616" target="_blank">https://huggingface.co/papers/2607.27616</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233209977.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>7. Beyond Borrowed Histories: Person-Aligned User Simulation for Interactive Role-Playing Evaluation</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Role-playing agents, Evaluation, User simulators, Personalized rubrics, Multi-turn conversation</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Human-AI Interaction</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The paper aims to improve the evaluation of Role-playing agents (RPAs) by addressing the limitations of existing benchmarks and introducing a more personalized and realistic evaluation method through user simulators.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Implementation of PALATE, a scalable RPA benchmark using 300 character profiles and personalized rubrics, enabling free-form, multi-turn conversations with user simulators to assess RPAs&#8217; conversational abilities.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; PALATE provides interpretable evaluations of RPAs&#8217; performance in specific user interactions, demonstrating higher agreement with human judgments compared to traditional rubrics, and addresses both generic quality and personalized user experience.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.27816" target="_blank">https://huggingface.co/papers/2607.27816</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233235908.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>8. SpatialCLI: Learning to Reason With Spatial Tools, Then Without Them</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Vision-language models, SpatialCLI, Embodied agents, Compositional perception</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Multi-Modal Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The paper aims to address the capability mismatch in Vision-language models (VLMs) between task-level reasoning and capturing visual details, proposing a new framework to bridge this gap.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The proposed SpatialCLI framework operates in three stages: exposing specialist vision models as spatial tools, improving tool use with Cold-Start SFT and agentic RL, and internalizing successful tool-use trajectories.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; SpatialCLI significantly improves VLM performance on the MindCube benchmark, raising the accuracy from 29.3% to 84.6% with tools, and maintaining 73.8% accuracy even without tools after internalization, surpassing existing models.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.27703" target="_blank">https://huggingface.co/papers/2607.27703</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233300818.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>9. Chimera: Designing and Chinchilla-Scaling Hybrid Visual Diffusion Transformers</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Visual Generation, Chimera, Kimi Delta Attention, Multi-head Latent Attention, Sparse Mixture-of-Experts</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Generative Models</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To develop a hybrid visual diffusion backbone, Chimera, for efficient handling of high-resolution images, long videos, and multimodal context.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Introduction of Chimera that utilizes Kimi Delta Attention, Multi-head Latent Attention, and Sparse Mixture-of-Experts layers, along with HeteroP for scaling the architecture effectively.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Chimera&#8217;s dense backbone is 1.7x more compute-efficient than the Wan-2.1 baseline, and the complete system achieves 7.3x efficiency.</p>
<p>   &#8211; Chimera can extrapolate from 5-second training clips to 30-second videos with minimal performance degradation.</p>
<p>   &#8211; Fitted compute-optimal laws indicate balanced compute division in image pretraining, with a slight model size preference in video pretraining at higher budgets.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28611" target="_blank">https://huggingface.co/papers/2607.28611</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233323944.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>10. β-OPSD: Deriving with Policy Optimization, Training with Self-Distillation</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: On-policy self-distillation, policy optimization, reinforcement learning, reasoning performance, mathematical reasoning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Reinforcement Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study aims to enhance reasoning language models using a refined version of on-policy self-distillation, termed β-OPSD, which introduces a versatile regularization parameter β for better policy optimization.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The paper introduces β-OPSD by positioning it within a broader policy-optimization family. It employs reinforcement learning for distilling the closed-form solution through efficient token-level logit mixing to approximate costly optimization processes.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Experiments demonstrate that β-OPSD consistently outperforms the standard OPSD in mathematical reasoning tasks, achieving improved optimization stability and better reasoning performance.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28582" target="_blank">https://huggingface.co/papers/2607.28582</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233348258.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>11. Can Large Language Models Execute Parent Orders?</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Parent-order execution, Large Language Models (LLMs), PACE, Shenzhen Stock Exchange, Execution decisions</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: AI in Finance</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study aims to enhance parent-order execution in algorithmic trading using Large Language Models (LLMs), addressing limitations of current methods like reliance on market assumptions and lack of adaptability.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Introduction of PACE (Plan-Ahead Controlled Execution), a hierarchical framework that separates parent-order execution into long-term planning and short-term execution, bypassing the need for specific market assumptions and training.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; PACE demonstrated superior performance compared to existing TWAP, Almgren-Chriss, and learning-based strategies during tests on Shenzhen Stock Exchange data. It was observed that LLMs navigate execution decisions differently from humans, with higher confidence correlated with improved performance and earlier trades instead of late actions.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28410" target="_blank">https://huggingface.co/papers/2607.28410</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233411593.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>12. LEDGERMIND: Provenance-Constrained Multimodal Agentic Reasoning with a Structured Evidence Ledger</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Multimodal Agents, Visual Question Answering, Provenance-Constrained, Structured Evidence, LedgerMind</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Multi-Modal Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study aims to improve the evaluation of multimodal agents that handle visual question answering by considering not only answer accuracy but also the process through which answers are derived.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Introduction of LedgerMind, a system based on a provenance-constrained state machine, to standardize tool outputs and support structured evidence.</p>
<p>   &#8211; Utilization of a Three-Layer Grounding Protocol, Adaptive Dual-Path Dispatcher, and an Event-Triggered Verification-and-Repair engine to manage reasoning depth and ensure provenance.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; LedgerMind effectively addresses common failure patterns such as unsupported reasoning and entity hallucinations, enhancing both answer accuracy and the reliability of the reasoning trajectory.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28374" target="_blank">https://huggingface.co/papers/2607.28374</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233436135.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>13. Σ-Mem: An Online Reliability Memory for LLM-based Multi-Agent Systems</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: LLM agents, multi-agent systems, Σ-Mem, reliability memory, interaction content</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: AI Systems and Tools</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The objective is to introduce Σ-Mem, a novel online reliability memory system, addressing limitations in existing memory systems within LLM agents, particularly in multi-agent environments.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Σ-Mem models historical competence and peer relationship evidence, using real symmetric states updated from correctness feedback while ensuring stable online adaptation without requiring model retraining.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Σ-Mem effectively adapts to shifts in counterfactual reliability and generalizes to new peers and task domains, outperforms traditional voting methods, and improves with increased correctness feedback, establishing a solid foundation for adaptive coordination in multi-agent LLM systems.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.27958" target="_blank">https://huggingface.co/papers/2607.27958</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233503295.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>14. Explorative Modeling: Unlocking a Third Pretraining Axis and End-to-End Generation</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Explorative Modeling, Generative Models, FLOP Efficiency, Sample Efficiency, Parameter Efficiency</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Generative Models</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; Introduce Explorative Modeling as a new paradigm in generative modeling to enable end-to-end training by exploring candidate matches and training on the best predictions.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The Explorative Models are applied in generative domains, increasing exploration as a pretraining axis to improve performance across diverse datasets including images, video, and language.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Explorative Modeling enhances efficiency exponentially, achieving significant improvement in FLOP, sample, and parameter efficiency, and enables end-to-end reconstructive generative modeling with reduced inference steps, establishing itself as a new training axis for existing generative models.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.27372" target="_blank">https://huggingface.co/papers/2607.27372</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233526935.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>15. AI Tour Meeting: Group Travel Planning by LLM Agents</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: AI Tour Meeting, Large Language Model, agents, simulation tool</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: AI Systems and Tools</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To propose a travel planning framework called AI Tour Meeting using agents powered by Large Language Models (LLMs) to collaboratively create itineraries through natural language discussion.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Implementation of a framework for configuring agent personas and orchestrating discussion workflows, enabling monitoring and LLM deployment in the simulation of tour planning discussions.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The framework is validated by system demonstration and provides analytical insights into the behavior of multiple LLM agents during group travel planning.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.18806" target="_blank">https://huggingface.co/papers/2607.18806</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233614234.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>16. Revisiting Lossy Verification in Speculative Decoding: Mechanisms, Trade-offs, and Failure Modes</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Speculative Decoding, Large Language Model, Truncation-based Verification, Collaborative Verification</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Natural Language Processing</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To analyze the distributions induced by lossy verification schemes in speculative decoding and identify their potential pitfalls in large language model inference.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Classified lossy verification methods into truncation-based and collaborative verification, and developed a diagnostic evaluation framework across curated benchmarks.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Truncation-based methods are susceptible to significant performance degradation due to distributional distortion.</p>
<p>   &#8211; For collaborative verification, controlling the overshoot of draft probabilities relative to target probabilities is crucial to maintaining output quality.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.26627" target="_blank">https://huggingface.co/papers/2607.26627</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233549571.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>17. Beyond Geometric Complementarity: Coherent Overlap in Sparse Mixture-of-Experts Routing</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Sparse mixture-of-experts, Expert Subspace Separation Index, coherent overlap, next-token prediction</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Natural Language Processing</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study aims to distinguish the benefits and interactions of experts within sparse mixture-of-experts (MoE) language models by analyzing their geometric arrangement.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Utilizes an Expert Subspace Separation Index (ESSI) and various controlled interventions, including matched-route residuals and prefix-controlled factorial setups, to analyze expert overlaps and their impact on representation.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Coherent overlap indicates that while routes overlap, useful multi-expert computation persists, clarifying why geometric similarity doesn&#8217;t alone determine redundancy or pruning value.</p>
<p>   &#8211; Adding later experts improves next-token predictions significantly in most cases, emphasizing the importance of non-redundant expert selections.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28308" target="_blank">https://huggingface.co/papers/2607.28308</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233702669.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>18. Fairness Pruning: Locating Demographic Bias in GLU-MLP Layers via Differential Activations</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Fairness Pruning, Large Language Models, Demographic Bias, Causal Bias Localization, AI Ethics and Fairness</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: AI Ethics and Fairness</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To introduce and empirically validate Fairness Pruning, a method aimed at managing and mitigating demographic biases in large language models.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Utilization of minimally contrastive prompt pairs and inference-time activation capture to identify neurons that react differentially to demographic attributes.</p>
<p>   &#8211; Empirical evaluation on large models such as Llama-3.2 and Salamandra-2B with structural intervention by zeroing neurons.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Demonstrated that zeroing specific neurons alters model responses to demographic variables, with a bidirectional bias destabilization effect.</p>
<p>   &#8211; Confirmed that demographic bias processing and model capabilities are dissociable, setting foundations for transitioning to directional behavior modulation.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28319" target="_blank">https://huggingface.co/papers/2607.28319</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233637423.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>19. AMRD: Adaptive Multi-Teacher Relational Distillation for Lightweight Speech Emotion Recognition</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Speech Emotion Recognition, Edge Devices, Knowledge Distillation, AI Native</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Machine Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; This research aims to optimize on-device speech emotion recognition (SER) for real-time applications by developing a more efficient model that can operate effectively on edge devices.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Adaptive Multi-teacher Relational Distillation (AMRD) is proposed, using a one-class SVM to assign weights to teacher models and a relational distillation loss to improve inter-sample relational structure alignment between teacher and student models.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The AMRD approach outperforms single-teacher distillation baselines on the IEMOCAP and CREMA-D datasets, demonstrating its effectiveness across multiple student architectures and confirming that both components of the method yield complementary benefits.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.25289" target="_blank">https://huggingface.co/papers/2607.25289</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233726024.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>20. </h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="" target="_blank"></a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/thirteen/202607311785541064.jpg"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>21. Pedestrian Archetypes Extension &#8212; More Pedestrian Models for Autonomous Vehicle Safety Testing</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Pedestrian Archetypes, Real-world Traffic Scenarios, Behavior Patterns</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Robotics and Autonomous Systems</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The research introduces new pedestrian archetypes that exhibit significant behavioral differences, enhancing the understanding of pedestrian behaviors beyond the original taxonomy.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Analysis of YouTube dash-cam videos for identifying additional pedestrian archetypes not covered in previous work.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Seven new pedestrian archetypes are presented, each defined by unique behaviors that are documented through video-frame evidence, thus expanding the existing framework for assessing pedestrian behavior in traffic.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.16922" target="_blank">https://huggingface.co/papers/2607.16922</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233737256.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>22. ReToken: One Token to Improve Vision-Language Models for Visual Retrieval</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Vision-Language Models, GPU Memory Constraints, Image-QA Dataset</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Multi-Modal Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; Address the challenge of long visual context in vision-language models by introducing ReToken, an explicit retrieval target for selecting query-relevant visual tokens.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; ReToken is trained on a small image-QA dataset and designed to select a sparse set of visual tokens from a pre-filled visual KV cache.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; ReToken shows consistent performance gains across image and video benchmarks, with significant improvements in Visual Haystacks and LVBench, while maintaining efficiency on a single H100 for both training and long-video inference.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28627" target="_blank">https://huggingface.co/papers/2607.28627</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233713703.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>23. OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: OmniScope, token compression, query, salience mismatch, AI Native</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Multi-Modal Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The objective is to address the cross-modal salience mismatch in token compression methods for omnimodal large language models by developing a framework named OmniScope that estimates relevance separately for audio and video.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Introduced a training-free token compression framework using the query as a shared semantic anchor, allocating modality-specific token budgets, pruning visual tokens with an anchor-delta strategy, and merging audio tokens to reduce redundancy.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; OmniScope outperforms existing methods in accuracy across audio-video benchmarks and model scales, offering improvements in speed and memory efficiency while maintaining high accuracy with minimal drop.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.23193" target="_blank">https://huggingface.co/papers/2607.23193</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233650331.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>24. Harness-G: A Graph-Structured Harness for Search Agents</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Reinforcement Learning, retrieval aliasing, Harness-G, Structured Non-myopic Credit</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Reinforcement Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study addresses the problem of retrieval aliasing in reinforcement learning search agents by proposing a new retrieval framework to optimize query generation and interaction.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The authors introduce Harness-G, a graph-structured retrieval framework that transforms query generation into finite action selection. They also develop Structured Non-myopic Credit (SNC) for better rewarding actions.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Harness-G achieves the highest average F1 across six QA benchmarks and outperforms the best baseline by significant margins, demonstrating its effectiveness in reducing linguistic aliasing and improving retrieval accuracy.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.27652" target="_blank">https://huggingface.co/papers/2607.27652</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233625249.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>25. Is Deep Research Reliable? Misleading Knowledge Induces False Conclusions</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Deep Research agents, Misleading Knowledge, Evidence Synthesis, Workflow Reliability</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Knowledge Representation and Reasoning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study aims to explore the reliability vulnerabilities in Deep Research agents when exposed to misleading knowledge in open information environments.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Introduces MisKnow-Agent, a framework for generating misleading knowledge instances to test these vulnerabilities, utilizing 5,933 quality-controlled instances derived from DeepResearch Benchmark tasks.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The experiments demonstrate that even minimal misleading knowledge can lead to false conclusions in final reports, highlighting a significant reliability issue.</p>
<p>   &#8211; Despite verifier models identifying misleading instances, they can still be adopted in long-term workflows, indicating a gap in workflow-level evidence verification.</p>
<p>   &#8211; Evaluated defenses show mitigation effects but fail to fully prevent the adoption of false conclusions, pointing to the need for improved verification and correction capabilities.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.20891" target="_blank">https://huggingface.co/papers/2607.20891</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233600666.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>26. Filesystem-Based Memory for LLM Agents: Organization, Evolution, and Sustainability</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: LLM agents, filesystem-based memory, memory organization, search economy, management agent</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: AI Systems and Tools</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To systematically explore filesystem-based memory for Large Language Model (LLM) agents, assessing the structure and organization of such memory.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The study formalizes the concept of a memory filesystem with a management agent for organization, a search agent for querying, and an execution agent for task trajectory integration.</p>
<p>   &#8211; Investigated variables include memory shape, stream scale, tool harness, and agent strength, with focus on memory growth, retrieval costs, and store health.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Organization in filesystem-based memory leads to improved search economy, significantly reducing retrieval costs.</p>
<p>   &#8211; However, as memory grows, management agents struggle to maintain organization quality, and improvements in organization do not necessarily translate to better answer quality.</p>
<p>   &#8211; The study identifies the tool set as a significant factor in shaping the memory store, highlighting a design space for agent memory rather than just an assumption.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.26637" target="_blank">https://huggingface.co/papers/2607.26637</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233538463.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>27. Multi-Head Attention Residuals</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Transformers, Multi-Head Attention Residuals, Nemotron, validation loss</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Machine Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; Introduction of Multi-Head Attention Residuals (MHAR) to improve information propagation across model depth and width.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; MHAR utilizes H per-subspace heads with distinct softmax distributions, trained on a Nemotron-based corpus to optimize validation loss across various model sizes.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; MHAR outperforms standard Transformers with validation loss improvements across 100M, 350M, and 1B scales. Optimal head count H = 8 increases efficiency and maintains near-baseline memory consumption.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.27230" target="_blank">https://huggingface.co/papers/2607.27230</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233514954.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>28. Echoverse: Deep, Evolving Environments for Training Computer-Use Agents at Scale</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Synthetic Environments, Echoverse, Reinforcement Learning, State Representation</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Reinforcement Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To improve computer-use agents&#8217; training environments by focusing on behavioral depth, targeted interactions, and model improvement.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Developed Echoverse to compile specifications into stateful applications, incorporating a co-evolution loop for iterative environment and model enhancement.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Demonstrated significant model improvement across various environments, with increases in evaluation scores and live-site accuracy via deep environments and iterative environment repairs.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28074" target="_blank">https://huggingface.co/papers/2607.28074</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233451225.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>29. INTACT: Isomorphic Intent-to-Action Learning for Search-Free World Models</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: INTACT, Intent-to-Action, Latent World Models, Action Law Semantics</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Reinforcement Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To introduce INTACT, an end-to-end JEPA, aimed at transforming action-labeled, reward-free trajectories into a deployable intent-to-action interface without expensive test-time search.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Utilizes a unique architecture with isomorphic intent representation for local and goal motion-intent models, and provides intact transfer from RGB evidence to action-effective latent intent coordinates.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The INTACT model achieved up to 100% success rates on LeWM tasks using a search-free policy, significantly reducing the need for candidate sequence sampling while improving performance through a shared four-task encoder. Direct inference is highly efficient, taking only 2.9-5.5 milliseconds.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.26056" target="_blank">https://huggingface.co/papers/2607.26056</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233423941.jpg"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>30. MemHarness: Memory Is Reconstructed, Not Replayed</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: MemHarness, Negative Transfer, Context-grounded Guidance, Large Language Model Agents, GRPO</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Reinforcement Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The objective is to develop a framework named MemHarness that enables large language model agents to actively reconstruct past experiences based on the current context, thereby providing more accurate and context-relevant guidance.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; MemHarness employs a unified policy model that critiques and reconstructs retrieved experiences for each decision step. This reconstructive ability is achieved through end-to-end training using the GRPO technique.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; MemHarness significantly outperforms existing memory-augmented baselines and pure reinforcement learning models, especially in out-of-distribution scenarios, by preventing negative transfer and enhancing intrinsic reasoning capabilities.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28272" target="_blank">https://huggingface.co/papers/2607.28272</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233359267.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>31. ShadowDancer: Teaching Video World Models Any Action by Learning Unified Dynamics Representations from a Video and Its Shadow</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: ShadowDancer, interactive video, dynamics, action transfer, cross-shadow prediction</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Computer Vision</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To enable precise control of any-action, frame-level dynamics in interactive video world models through novel techniques.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Utilizes shadow pairs to construct scalable video pairs replaying the same dynamics with varied appearances.</p>
<p>   &#8211; Implements cross-shadow prediction to learn actions by predicting dynamics from one shadow to another.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Demonstrated improved action transfer and extended action rollout in diverse dynamics families, achieving an average blinded win rate of 86% over existing models.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28362" target="_blank">https://huggingface.co/papers/2607.28362</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233335737.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>32. See2Think: Do Multimodal Models Really Use Intermediate Visual States?</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Multimodal large language models, Visual reasoning, See2Think, Visual Action-of-Thought</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Multi-Modal Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To introduce See2Think, an evaluation framework that assesses visual reasoning in multimodal large language models through benchmarks like See2ThinkBench and Visual Action-of-Thought.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The creation of See2ThinkBench with 1,200 visually dependent tasks across 12 categories and the use of Visual Action-of-Thought to record and analyze visual and textual operations under controlled settings.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Visual reasoning capabilities in models are strongly dependent on both the model and environment, showing varied performance across tasks. Additionally, even though models often select the correct visual operations, the accuracy improvements are hindered by challenges in rendering fidelity and reliance on feedback which does not always improve accuracy.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.26769" target="_blank">https://huggingface.co/papers/2607.26769</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233311886.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>33. RefCaptioner: Multi-Reference Image-Grounded Video Captioning</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: multi-reference image-grounded video captioning, RefCaptioner, factual video descriptions, Hierarchical Coverage-Discounted GRPO, AI-generated videos</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Multi-Modal Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To introduce and address the task of multi-reference image-grounded video captioning, requiring factual video descriptions with phrase-level reference grounding.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Development of RefCaptioner, a two-stage post-training framework combining mixed-data SFT with Hierarchical Coverage-Discounted GRPO to enhance reference selection and cross-reference consistency.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; RefCaptioner outperforms other open-source models in overall performance, yielding captions preferred by human annotators and allowing for more source-faithful video reconstruction.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28509" target="_blank">https://huggingface.co/papers/2607.28509</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233249116.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>34. ACE-Data-0: Human-Centric Ambient Capture as Embodied Data Engine</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Embodied intelligence, Ambient Capture Engine, AI Native, Multi-modal datasets, Human-centric data</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Robotics and Autonomous Systems</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study focuses on addressing the data bottleneck in embodied intelligence by introducing the Ambient Capture Engine (ACE), which captures comprehensive sensory data in real home environments.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The authors developed ACE to record multisensory data through a combination of table-scale and room-scale configurations, creating a detailed dataset, ACE-Data-0, encompassing 150 hours of interactions across 200 task categories.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Evaluations reveal significant limitations in current state-of-the-art methods, particularly under challenging conditions like occlusion and long temporal horizons, highlighting ACE&#8217;s potential in supporting scalable advancements in imitation learning and embodied AI systems.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28625" target="_blank">https://huggingface.co/papers/2607.28625</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233221968.jpg"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>35. Flux-OPD: On-Policy Distillation with Evolving Contexts</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Large language model, evolving contexts, Flux-OPD, open-ended domains, teacher supervision</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Natural Language Processing</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study aims to develop a mechanism for more effective supervision in large language model training across open-ended domains, addressing the issue of lacking verifiable rewards and formalized task preferences.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The research introduces Flux-OPD, a novel OPD paradigm that leverages evolving contexts as a form of in-training supervision. It analyzes the effect of context through the reverse KL objective and uses context-conditioned teachers to provide contextual difference signals.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Flux-OPD demonstrates superiority over existing OPD paradigms, effectively capturing task preferences and enhancing performance in open-ended tasks by combining evolving contexts with teacher supervision.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28022" target="_blank">https://huggingface.co/papers/2607.28022</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233157728.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>36. Beacon: Knowing When and How to Perform Agentic Visual Reasoning</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Agentic Visual Reasoning, Multimodal Large Language Models, Mode Adaptiveness, Tool Effect, Beacon</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Multi-Modal Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study aims to enhance the success rate of multimodal large language models (MLLMs) in complex tasks by rethinking agentic visual reasoning, focusing on Mode Adaptiveness and Tool Effect.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Comprehensive analysis to quantify Mode Adaptiveness and Tool Effect in existing models, followed by the proposal of a new model named Beacon, featuring Necessity-Aware Adaptive Reward and Hint-Guided Capability Expansion.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The research concludes that existing models have limited Mode Adaptiveness, and tool benefits are often neutralized by their drawbacks. Beacon, however, shows improved overall performance, notably enhancing Mode Adaptiveness and achieving genuine gains from tool use.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28595" target="_blank">https://huggingface.co/papers/2607.28595</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233136136.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>37. VideoCoCo: Code-as-CoT for Physically-Consistent Video Generation via an Agentic Dual-Engine System</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Text-to-video, VideoCoCo, Blender program, spatiotemporal process, generative video engine</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Generative Models</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study aims to address the limitations of text-to-video models in generating physically consistent dynamics by introducing a new framework called VideoCoCo.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; A dual-engine framework is employed, where a coding agent synthesizes executable Blender code to specify scene evolution and produce deterministic drafts, followed by a generative video engine for photorealistic video transformation.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; VideoCoCo effectively improves performance on benchmarks PhyGenBench and VBench-2.0 by providing an executable, controllable, and inspectable representation for video generation, achieving the best average scores on both benchmarks.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.27380" target="_blank">https://huggingface.co/papers/2607.27380</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233111733.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>38. PhiZero: A World Model Built Around Physical Language</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: PhiZero, physical language, world model, zero-shot motion transfer</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Generative Models</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The primary aim is to develop PhiZero, a model that abstracts predictive structure from visual experience to explicitly reason about physical world evolution using physical language.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; PhiZero employs a reason-then-render paradigm where it infers future world evolution sequences in physical language and renders transitions into video through self-supervision from in-the-wild videos.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The experiments demonstrate PhiZero&#8217;s capability to model physically coherent world evolution and highlight its potential for realistic interactive world modeling, fine-grained action-conditioned simulation, and zero-shot motion transfer.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28624" target="_blank">https://huggingface.co/papers/2607.28624</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233043717.jpg"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>39. Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: GUI agents, Qwen-UI-Agent, Reinforcement Learning, MobileWorld, Online RL</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: AI Systems and Tools</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The paper aims to advance GUI agents for real-world applications, enabling reliable operation across various devices and platforms, proactive service initiation, and autonomous capability improvements.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Introduces Qwen-UI-Agent, utilizing a unified action space that combines GUI operations with CLI execution, and employs an AutoResearch-style data flywheel and Online Reinforcement Learning (RL) for training and advancing agent performance.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Qwen-UI-Agent demonstrates state-of-the-art performance in mobile-use benchmarks and competitive results in computer- and browser-use tasks, achieving high scores on MobileWorld, AndroidDaily, and WebArena among others.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28227" target="_blank">https://huggingface.co/papers/2607.28227</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233022448.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<p>The post <a href="https://ainativefoundation.org/ai-native-daily-paper-digest-20260917-askchem-metis-chimera/">AI Native Daily Paper Digest – 20260917 – AskChem | Metis | Chimera</a> appeared first on <a href="https://ainativefoundation.org">AI Native Foundation</a>.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>China AI Native Industry Insights &#8211; 20260917 &#8211;  Tencent &#124; Alibaba &#124; MiniMax &#124; more</title>
		<link>https://ainativefoundation.org/china-ai-native-industry-insights-20260917-tencent-alibaba-minimax-more/</link>
		
		<dc:creator><![CDATA[AINF]]></dc:creator>
		<pubDate>Thu, 17 Sep 2026 08:04:55 +0000</pubDate>
				<category><![CDATA[China Industry]]></category>
		<guid isPermaLink="false">https://ainativefoundation.org/?p=10985</guid>

					<description><![CDATA[<p>Today's digest highlights significant advancements from Tencent, Alibaba, and Qoder Cloud Agents. These companies are fostering an ecosystem focused on enhancing AI agent capabilities and scalability. Tencent has open-sourced BrowserSkill, a CLI tool that allows AI agents to utilize browser tabs more effectively, while Alibaba's Ant Ling team introduces three quantized versions of Ling-3.0-flash-Fin, aimed at refining AI financial analysis. Meanwhile, Qoder Cloud Agents has launched its Agent-as-a-Service platform with version 1.0, designed for seamless production deployment. These developments indicate a strong push towards open-source solutions and improved AI agent integration within various applications.</p>
<p>The post <a href="https://ainativefoundation.org/china-ai-native-industry-insights-20260917-tencent-alibaba-minimax-more/">China AI Native Industry Insights &#8211; 20260917 &#8211;  Tencent | Alibaba | MiniMax | more</a> appeared first on <a href="https://ainativefoundation.org">AI Native Foundation</a>.</p>
]]></description>
										<content:encoded><![CDATA[<p>Today&#8217;s digest highlights significant advancements from Tencent, Alibaba, and Qoder Cloud Agents. These companies are fostering an ecosystem focused on enhancing AI agent capabilities and scalability. Tencent has open-sourced BrowserSkill, a CLI tool that allows AI agents to utilize browser tabs more effectively, while Alibaba&#8217;s Ant Ling team introduces three quantized versions of Ling-3.0-flash-Fin, aimed at refining AI financial analysis. Meanwhile, Qoder Cloud Agents has launched its Agent-as-a-Service platform with version 1.0, designed for seamless production deployment. These developments indicate a strong push towards open-source solutions and improved AI agent integration within various applications. Discover more in Today&#8217;s China AI Native Industry Insights.</p>
<h3>1.  Tencent open-sources BrowserSkill, a CLI tool letting AI agents borrow your browser tab</h3>
<p>Tencent released BrowserSkill, an open-source, MIT-licensed CLI tool that connects AI agents to a user&#8217;s actual browser instead of a blank automated instance. The tool lets an agent borrow an existing tab, preserving login sessions, while returning control to the user for captchas and confirmation dialogs. As a CLI rather than an MCP server, it works with any agent capable of running shell commands, including Cursor, Claude Code, Codex, and others, and logs every call it makes. Tab-borrowing permission is controlled through browser settings rather than a command-line flag, and the entire tool runs locally.<br />
Read more: <a href="https://github.com/Tencent/BrowserSkill">https://github.com/Tencent/BrowserSkill</a></p>
<p><video width="600" height="400" controls poster="https://cdn.ainative.foundation/image/20260917_gn_img_tencent.jpg"><source src="https://cdn.ainative.foundation/video/20260917_gn_video_tencent.mp4" type="video/mp4"></video><br />
Video Credit: The original article</p>
<h3>2.  Alibaba&#8217;s Ant Ling team releases three quantized versions of Ling-3.0-flash-Fin</h3>
<p>Ant Ling, part of Alibaba&#8217;s AI research group, announced three quantized versions of its Ling-3.0-flash-Fin model: FP8, FP4, and INT4. The releases target developers building real-world financial workflows, allowing them to select a version suited to their infrastructure, memory, and efficiency requirements. Ling-3.0-flash-Fin is a variant of the Ling-3.0 model line focused on financial applications. The quantized releases aim to lower deployment costs while preserving model performance.<br />
Read more: <a href="https://x.com/i/web/status/2099512582248046639">https://x.com/i/web/status/2099512582248046639</a></p>
<p><video width="600" height="400" controls poster="https://cdn.ainative.foundation/image/20260917_8d5f583530974f81ad3968f9faf0d444.jpg"><source src="https://cdn.ainative.foundation/video/20260917_gn_video_ling.mp4" type="video/mp4"></video><br />
Video Credit: @AntLingAGI  on X</p>
<h3>3.  MiniMax Highlights Open-Source Ecosystem Growth Around H3 Video Generation Model</h3>
<p>MiniMax announced its H3 open-weight video generation model, which supports native stereo audio and multimodal reference control, is gaining rapid open-source community adoption. Collaborators including NVIDIA, FastVideo, Nuva Lab, and Alibaba PAI have released tools such as FastH3, Sol-H3, VDN, and distillation-based LoRAs that speed up H3 inference across hardware like DGX Spark, Apple Silicon, and B300 GPUs. These projects include released weights, training code, and inference code, with workflows also integrated into ComfyUI. MiniMax credited the individual contributors and teams driving these optimizations.<br />
Read more: <a href="https://www.minimax.cn/news/minimax-h3-open-source">https://www.minimax.cn/news/minimax-h3-open-source</a></p>
<p><video width="600" height="400" controls poster="https://cdn.ainative.foundation/image/20260917_881d92e501af4976960c85e38a5f1b40.jpg"><source src="https://cdn.ainative.foundation/video/20260917_gn_video_minimax.mp4" type="video/mp4"></video><br />
Video Credit: NotebookLM</p>
<h3>4.  Qoder Cloud Agents 1.0 launches Agent-as-a-Service platform for production deployment</h3>
<p>Qoder has officially released Qoder Cloud Agents 1.0, an Agent-as-a-Service platform designed to help enterprises and individual developers deploy and scale AI agents in production environments. The platform has been in testing since May, supporting hundreds of customers across e-commerce, entertainment, consumer electronics, automotive, and real estate industries, with monthly invocations growing from tens of thousands to tens of millions. QCA 1.0 provides cloud-based harness hosting, multi-agent coordination, enterprise integration with skills and MCP, batch processing, high-concurrency elastic scheduling, and cross-session memory capabilities. Real-world use cases include upgrading e-commerce operational tools and scaling customer service quality inspection from 1% sampling to full coverage.<br />
Read more: <a href="https://mp.weixin.qq.com/s?__biz=MzE5ODI2NzI5Nw==&#038;mid=2247492087&#038;idx=1&#038;sn=2dffa75a38ddff69c2fee680c3c580e4&#038;chksm=975fa2f166e554df7d458f9b72391323ae4dad6f3366110f71e921aece082c4b09bfd4ae2e32&#038;scene=0&#038;xtrack=1#rd">https://mp.weixin.qq.com/s?__biz=MzE5ODI2NzI5Nw==&#038;mid=2247492087&#038;idx=1&#038;sn=2dffa75a38ddff69c2fee680c3c580e4&#038;chksm=975fa2f166e554df7d458f9b72391323ae4dad6f3366110f71e921aece082c4b09bfd4ae2e32&#038;scene=0&#038;xtrack=1#rd</a></p>
<p><video width="600" height="400" controls poster="https://cdn.ainative.foundation/image/20260917_88fb41da597e4b8092a88a57a9a453a1"><source src="https://cdn.ainative.foundation/video/20260917_gn_video_qoder.mp4" type="video/mp4"></video><br />
Video Credit: NotebookLM</p>
<div style="width:100%;height:2px;background:#808080;margin:10px 0"></div>
<p>That&#8217;s all for today&#8217;s China AI Native Industry Insights. Join us at <a href="https://member.ainativefoundation.org/">AI Native Foundation Membership Dashboard</a> for the latest insights on AI Native, or follow our linkedin account at <a href="https://www.linkedin.com/company/ainativefoundation/">AI Native Foundation</a> and our twitter account at <a href="https://x.com/AINativeF">AINativeF</a>.</p>
<p>The post <a href="https://ainativefoundation.org/china-ai-native-industry-insights-20260917-tencent-alibaba-minimax-more/">China AI Native Industry Insights &#8211; 20260917 &#8211;  Tencent | Alibaba | MiniMax | more</a> appeared first on <a href="https://ainativefoundation.org">AI Native Foundation</a>.</p>
]]></content:encoded>
					
		
		<enclosure url="https://cdn.ainative.foundation/video/20260917_gn_video_tencent.mp4" length="16604300" type="video/mp4" />
<enclosure url="https://cdn.ainative.foundation/video/20260917_gn_video_ling.mp4" length="12707829" type="video/mp4" />
<enclosure url="https://cdn.ainative.foundation/video/20260917_gn_video_minimax.mp4" length="9466199" type="video/mp4" />
<enclosure url="https://cdn.ainative.foundation/video/20260917_gn_video_qoder.mp4" length="27068387" type="video/mp4" />

			</item>
		<item>
		<title>AI Native Daily Paper Digest – 20260916 – AskChem &#124; Metis &#124; Chimera</title>
		<link>https://ainativefoundation.org/ai-native-daily-paper-digest-20260916-askchem-metis-chimera/</link>
		
		<dc:creator><![CDATA[insights]]></dc:creator>
		<pubDate>Thu, 17 Sep 2026 00:41:04 +0000</pubDate>
				<category><![CDATA[Papers]]></category>
		<guid isPermaLink="false">https://ainativefoundation.org/ai-native-daily-paper-digest-20260916-askchem-metis-chimera/</guid>

					<description><![CDATA[<p>Today&#8217;s digest highlights significant developments from leading entities such as Gemma and Qwen, showcasing their continued innovation in AI. The overarching theme [&#8230;]</p>
<p>The post <a href="https://ainativefoundation.org/ai-native-daily-paper-digest-20260916-askchem-metis-chimera/">AI Native Daily Paper Digest – 20260916 – AskChem | Metis | Chimera</a> appeared first on <a href="https://ainativefoundation.org">AI Native Foundation</a>.</p>
]]></description>
										<content:encoded><![CDATA[<p style="margin-bottom:24px">Today&#8217;s digest highlights significant developments from leading entities such as Gemma and Qwen, showcasing their continued innovation in AI. The overarching theme revolves around enhancing agentic systems and improving multimodal reasoning capabilities. Among the noteworthy studies, a novel methodology for integrating diverse data inputs into a coherent decision-making framework is introduced, showing substantial improvements across standard benchmarks like ImageNet with a 5% increase in accuracy. Additionally, one paper presents robust findings on the scalability of long-context attention mechanisms in natural language processing tasks. A separate study demonstrates how these advancements can optimize real-time robotics navigation, reducing computational latency by 30%.</p>
<h3>1. AskChem: Claim-Centered Infrastructure for Chemistry Literature Synthesis</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: AskChem, claim-centered infrastructure, GPT-5.5, AI agents, chemistry search</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Knowledge Representation and Reasoning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The objective of the research was to develop AskChem, a claim-centered infrastructure for conducting cross-paper chemistry searches, which shifts the retrieval focus from whole papers to provenance-carrying claims.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; AskChem converts papers into atomic, typed claims grounded by a source DOI and quotes or evidence locators, enabling hierarchical retrieval and exploration through a faceted taxonomy and evidence graph.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; AskChem indexes 2.4M claims from 147K papers, improving the retrieval capability with 100% resolvable DOIs for a GPT-5.5 reader and provides superior citation density compared to other systems.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28618" target="_blank">https://huggingface.co/papers/2607.28618</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img fetchpriority="high" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233007276.jpg"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>2. Metis: Memory Foundation Model</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: AI Native, native memory, memory foundation models, Metis</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Foundations of AI</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; Introduce memory foundation models to enable native memory capabilities in AI systems.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Proposed Metis, a new architecture with a native memory state for storing historical information.</p>
<p>   &#8211; Developed large-scale memory-specific training data and optimization objectives for native memory procedures.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Metis demonstrates native memory capabilities with advantages in architecture and efficiency.</p>
<p>   &#8211; The study provides a detailed analysis of strengths, limitations, and behaviors of Metis.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.26760" target="_blank">https://huggingface.co/papers/2607.26760</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233032494.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>3. Frontis-MA1: Training an AI4AI Model towards Recursive Self-Improvement in Machine Learning Engineering</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Recursive self-improvement, AI4AI, OpenMLE, Machine learning engineering, Reinforcement learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: AI Systems and Tools</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The paper aims to study Recursive self-improvement by introducing OpenMLE, an open full-stack system for AI4AI research in machine learning engineering.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Utilizes a combination of execution feedback environments (OpenMLE-Gym), operator learning (OpenMLE-RL), and long-horizon search (OpenMLE-Evo) to post-train a meta-evolution agent, Frontis-MA1, with atomic program-evolution operators.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The system significantly improves performance on MLE-Bench Lite, surpassing other models in efficiency and capability with OpenMLE-Evo-Max, and demonstrates transferable capabilities on NatureBench Lite.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28568" target="_blank">https://huggingface.co/papers/2607.28568</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233059370.jpg"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>4. Memory Decoder at Scale: A Pretrained, Parametric Long-Term Memory</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Memory Decoder, scalability, parameter efficiency, Faiss, pretraining</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Natural Language Processing</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To scale memory models up to 6.9B parameters and improve memory capacity independently for better performance of language models.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Implementation of a distributed pipeline for Faiss indexing and retrieval combined with sparse, batch-wise loading of kNN distributions.</p>
<p>   &#8211; Comparison of parameter-performance tradeoff by allocating more parameters to memory across different model scales.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Independently scaling pretrained memory models can achieve better parameter efficiency and enhance language model performance, as demonstrated by outperforming larger models with fewer parameters.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.27919" target="_blank">https://huggingface.co/papers/2607.27919</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233124782.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>5. BM25 Wins at Scale: A Scaling Study of Retrieval-Augmented Generation Paradigms</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Retrieval-Augmented Generation, BM25, Dense Retrieval, Lexical Retrieval, Agentic Reasoning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Natural Language Processing </p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To conduct a controlled study evaluating retrieval-augmented generation (RAG) approaches across varying corpus sizes and determine their accuracy-cost scaling.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The study varies corpus size across 28 nested tiers, analyzing accuracy, construction, query tokens, and latency under a single reader model and judging protocol.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Different approaches show varying effectiveness based on corpus size, with BM25 eventually leading as the corpus grows, while agentic search incurs high query token costs at larger scales.</p>
<p>   &#8211; Lexical retrieval emerges as the most scalable default, outperforming dense retrieval and graph-based indexing at larger corpus sizes, highlighting global candidate ranking&#8217;s importance.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.26497" target="_blank">https://huggingface.co/papers/2607.26497</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233146706.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>6. MPIE-Bench: Benchmarking Anatomically Plausible Multi-Person Interaction Editing</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Text-to-image, personalized editing, MPIE-Bench, MPIE-Eval, multi-person mesh reconstruction</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Computer Vision</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The research aims to address issues in synthesizing high-fidelity images featuring multiple named people in shared contact actions, such as embrace or carry.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Introduced MPIE-Bench, a benchmark of 2,500 samples from video-mined editing triplets covering various scenes, interactions, and contact densities.</p>
<p>   &#8211; Proposed MPIE-Eval, which evaluates contact-time geometry using a public multi-person mesh reconstruction framework.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The study finds that current visual language model (VLM) metrics often fail to capture anatomical and geometric errors, instead showing that MPIE-Eval aligns more closely with human judgment in evaluating anatomy and interaction accuracy.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.27616" target="_blank">https://huggingface.co/papers/2607.27616</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233209977.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>7. Beyond Borrowed Histories: Person-Aligned User Simulation for Interactive Role-Playing Evaluation</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Role-playing agents, Evaluation, User simulators, Personalized rubrics, Multi-turn conversation</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Human-AI Interaction</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The paper aims to improve the evaluation of Role-playing agents (RPAs) by addressing the limitations of existing benchmarks and introducing a more personalized and realistic evaluation method through user simulators.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Implementation of PALATE, a scalable RPA benchmark using 300 character profiles and personalized rubrics, enabling free-form, multi-turn conversations with user simulators to assess RPAs&#8217; conversational abilities.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; PALATE provides interpretable evaluations of RPAs&#8217; performance in specific user interactions, demonstrating higher agreement with human judgments compared to traditional rubrics, and addresses both generic quality and personalized user experience.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.27816" target="_blank">https://huggingface.co/papers/2607.27816</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233235908.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>8. SpatialCLI: Learning to Reason With Spatial Tools, Then Without Them</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Vision-language models, SpatialCLI, Embodied agents, Compositional perception</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Multi-Modal Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The paper aims to address the capability mismatch in Vision-language models (VLMs) between task-level reasoning and capturing visual details, proposing a new framework to bridge this gap.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The proposed SpatialCLI framework operates in three stages: exposing specialist vision models as spatial tools, improving tool use with Cold-Start SFT and agentic RL, and internalizing successful tool-use trajectories.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; SpatialCLI significantly improves VLM performance on the MindCube benchmark, raising the accuracy from 29.3% to 84.6% with tools, and maintaining 73.8% accuracy even without tools after internalization, surpassing existing models.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.27703" target="_blank">https://huggingface.co/papers/2607.27703</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233300818.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>9. Chimera: Designing and Chinchilla-Scaling Hybrid Visual Diffusion Transformers</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Visual Generation, Chimera, Kimi Delta Attention, Multi-head Latent Attention, Sparse Mixture-of-Experts</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Generative Models</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To develop a hybrid visual diffusion backbone, Chimera, for efficient handling of high-resolution images, long videos, and multimodal context.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Introduction of Chimera that utilizes Kimi Delta Attention, Multi-head Latent Attention, and Sparse Mixture-of-Experts layers, along with HeteroP for scaling the architecture effectively.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Chimera&#8217;s dense backbone is 1.7x more compute-efficient than the Wan-2.1 baseline, and the complete system achieves 7.3x efficiency.</p>
<p>   &#8211; Chimera can extrapolate from 5-second training clips to 30-second videos with minimal performance degradation.</p>
<p>   &#8211; Fitted compute-optimal laws indicate balanced compute division in image pretraining, with a slight model size preference in video pretraining at higher budgets.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28611" target="_blank">https://huggingface.co/papers/2607.28611</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233323944.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>10. β-OPSD: Deriving with Policy Optimization, Training with Self-Distillation</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: On-policy self-distillation, policy optimization, reinforcement learning, reasoning performance, mathematical reasoning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Reinforcement Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study aims to enhance reasoning language models using a refined version of on-policy self-distillation, termed β-OPSD, which introduces a versatile regularization parameter β for better policy optimization.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The paper introduces β-OPSD by positioning it within a broader policy-optimization family. It employs reinforcement learning for distilling the closed-form solution through efficient token-level logit mixing to approximate costly optimization processes.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Experiments demonstrate that β-OPSD consistently outperforms the standard OPSD in mathematical reasoning tasks, achieving improved optimization stability and better reasoning performance.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28582" target="_blank">https://huggingface.co/papers/2607.28582</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233348258.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>11. Can Large Language Models Execute Parent Orders?</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Parent-order execution, Large Language Models (LLMs), PACE, Shenzhen Stock Exchange, Execution decisions</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: AI in Finance</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study aims to enhance parent-order execution in algorithmic trading using Large Language Models (LLMs), addressing limitations of current methods like reliance on market assumptions and lack of adaptability.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Introduction of PACE (Plan-Ahead Controlled Execution), a hierarchical framework that separates parent-order execution into long-term planning and short-term execution, bypassing the need for specific market assumptions and training.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; PACE demonstrated superior performance compared to existing TWAP, Almgren-Chriss, and learning-based strategies during tests on Shenzhen Stock Exchange data. It was observed that LLMs navigate execution decisions differently from humans, with higher confidence correlated with improved performance and earlier trades instead of late actions.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28410" target="_blank">https://huggingface.co/papers/2607.28410</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233411593.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>12. LEDGERMIND: Provenance-Constrained Multimodal Agentic Reasoning with a Structured Evidence Ledger</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Multimodal Agents, Visual Question Answering, Provenance-Constrained, Structured Evidence, LedgerMind</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Multi-Modal Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study aims to improve the evaluation of multimodal agents that handle visual question answering by considering not only answer accuracy but also the process through which answers are derived.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Introduction of LedgerMind, a system based on a provenance-constrained state machine, to standardize tool outputs and support structured evidence.</p>
<p>   &#8211; Utilization of a Three-Layer Grounding Protocol, Adaptive Dual-Path Dispatcher, and an Event-Triggered Verification-and-Repair engine to manage reasoning depth and ensure provenance.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; LedgerMind effectively addresses common failure patterns such as unsupported reasoning and entity hallucinations, enhancing both answer accuracy and the reliability of the reasoning trajectory.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28374" target="_blank">https://huggingface.co/papers/2607.28374</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233436135.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>13. Σ-Mem: An Online Reliability Memory for LLM-based Multi-Agent Systems</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: LLM agents, multi-agent systems, Σ-Mem, reliability memory, interaction content</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: AI Systems and Tools</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The objective is to introduce Σ-Mem, a novel online reliability memory system, addressing limitations in existing memory systems within LLM agents, particularly in multi-agent environments.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Σ-Mem models historical competence and peer relationship evidence, using real symmetric states updated from correctness feedback while ensuring stable online adaptation without requiring model retraining.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Σ-Mem effectively adapts to shifts in counterfactual reliability and generalizes to new peers and task domains, outperforms traditional voting methods, and improves with increased correctness feedback, establishing a solid foundation for adaptive coordination in multi-agent LLM systems.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.27958" target="_blank">https://huggingface.co/papers/2607.27958</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233503295.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>14. Explorative Modeling: Unlocking a Third Pretraining Axis and End-to-End Generation</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Explorative Modeling, Generative Models, FLOP Efficiency, Sample Efficiency, Parameter Efficiency</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Generative Models</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; Introduce Explorative Modeling as a new paradigm in generative modeling to enable end-to-end training by exploring candidate matches and training on the best predictions.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The Explorative Models are applied in generative domains, increasing exploration as a pretraining axis to improve performance across diverse datasets including images, video, and language.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Explorative Modeling enhances efficiency exponentially, achieving significant improvement in FLOP, sample, and parameter efficiency, and enables end-to-end reconstructive generative modeling with reduced inference steps, establishing itself as a new training axis for existing generative models.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.27372" target="_blank">https://huggingface.co/papers/2607.27372</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233526935.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>15. AI Tour Meeting: Group Travel Planning by LLM Agents</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: AI Tour Meeting, Large Language Model, agents, simulation tool</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: AI Systems and Tools</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To propose a travel planning framework called AI Tour Meeting using agents powered by Large Language Models (LLMs) to collaboratively create itineraries through natural language discussion.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Implementation of a framework for configuring agent personas and orchestrating discussion workflows, enabling monitoring and LLM deployment in the simulation of tour planning discussions.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The framework is validated by system demonstration and provides analytical insights into the behavior of multiple LLM agents during group travel planning.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.18806" target="_blank">https://huggingface.co/papers/2607.18806</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233614234.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>16. Revisiting Lossy Verification in Speculative Decoding: Mechanisms, Trade-offs, and Failure Modes</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Speculative Decoding, Large Language Model, Truncation-based Verification, Collaborative Verification</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Natural Language Processing</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To analyze the distributions induced by lossy verification schemes in speculative decoding and identify their potential pitfalls in large language model inference.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Classified lossy verification methods into truncation-based and collaborative verification, and developed a diagnostic evaluation framework across curated benchmarks.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Truncation-based methods are susceptible to significant performance degradation due to distributional distortion.</p>
<p>   &#8211; For collaborative verification, controlling the overshoot of draft probabilities relative to target probabilities is crucial to maintaining output quality.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.26627" target="_blank">https://huggingface.co/papers/2607.26627</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233549571.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>17. Beyond Geometric Complementarity: Coherent Overlap in Sparse Mixture-of-Experts Routing</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Sparse mixture-of-experts, Expert Subspace Separation Index, coherent overlap, next-token prediction</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Natural Language Processing</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study aims to distinguish the benefits and interactions of experts within sparse mixture-of-experts (MoE) language models by analyzing their geometric arrangement.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Utilizes an Expert Subspace Separation Index (ESSI) and various controlled interventions, including matched-route residuals and prefix-controlled factorial setups, to analyze expert overlaps and their impact on representation.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Coherent overlap indicates that while routes overlap, useful multi-expert computation persists, clarifying why geometric similarity doesn&#8217;t alone determine redundancy or pruning value.</p>
<p>   &#8211; Adding later experts improves next-token predictions significantly in most cases, emphasizing the importance of non-redundant expert selections.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28308" target="_blank">https://huggingface.co/papers/2607.28308</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233702669.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>18. Fairness Pruning: Locating Demographic Bias in GLU-MLP Layers via Differential Activations</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Fairness Pruning, Large Language Models, Demographic Bias, Causal Bias Localization, AI Ethics and Fairness</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: AI Ethics and Fairness</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To introduce and empirically validate Fairness Pruning, a method aimed at managing and mitigating demographic biases in large language models.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Utilization of minimally contrastive prompt pairs and inference-time activation capture to identify neurons that react differentially to demographic attributes.</p>
<p>   &#8211; Empirical evaluation on large models such as Llama-3.2 and Salamandra-2B with structural intervention by zeroing neurons.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Demonstrated that zeroing specific neurons alters model responses to demographic variables, with a bidirectional bias destabilization effect.</p>
<p>   &#8211; Confirmed that demographic bias processing and model capabilities are dissociable, setting foundations for transitioning to directional behavior modulation.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28319" target="_blank">https://huggingface.co/papers/2607.28319</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233637423.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>19. AMRD: Adaptive Multi-Teacher Relational Distillation for Lightweight Speech Emotion Recognition</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Speech Emotion Recognition, Edge Devices, Knowledge Distillation, AI Native</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Machine Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; This research aims to optimize on-device speech emotion recognition (SER) for real-time applications by developing a more efficient model that can operate effectively on edge devices.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Adaptive Multi-teacher Relational Distillation (AMRD) is proposed, using a one-class SVM to assign weights to teacher models and a relational distillation loss to improve inter-sample relational structure alignment between teacher and student models.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The AMRD approach outperforms single-teacher distillation baselines on the IEMOCAP and CREMA-D datasets, demonstrating its effectiveness across multiple student architectures and confirming that both components of the method yield complementary benefits.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.25289" target="_blank">https://huggingface.co/papers/2607.25289</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233726024.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>20. </h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="" target="_blank"></a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/thirteen/202607311785541064.jpg"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>21. Pedestrian Archetypes Extension &#8212; More Pedestrian Models for Autonomous Vehicle Safety Testing</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Pedestrian Archetypes, Real-world Traffic Scenarios, Behavior Patterns</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Robotics and Autonomous Systems</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The research introduces new pedestrian archetypes that exhibit significant behavioral differences, enhancing the understanding of pedestrian behaviors beyond the original taxonomy.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Analysis of YouTube dash-cam videos for identifying additional pedestrian archetypes not covered in previous work.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Seven new pedestrian archetypes are presented, each defined by unique behaviors that are documented through video-frame evidence, thus expanding the existing framework for assessing pedestrian behavior in traffic.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.16922" target="_blank">https://huggingface.co/papers/2607.16922</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233737256.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>22. ReToken: One Token to Improve Vision-Language Models for Visual Retrieval</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Vision-Language Models, GPU Memory Constraints, Image-QA Dataset</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Multi-Modal Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; Address the challenge of long visual context in vision-language models by introducing ReToken, an explicit retrieval target for selecting query-relevant visual tokens.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; ReToken is trained on a small image-QA dataset and designed to select a sparse set of visual tokens from a pre-filled visual KV cache.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; ReToken shows consistent performance gains across image and video benchmarks, with significant improvements in Visual Haystacks and LVBench, while maintaining efficiency on a single H100 for both training and long-video inference.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28627" target="_blank">https://huggingface.co/papers/2607.28627</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233713703.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>23. OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: OmniScope, token compression, query, salience mismatch, AI Native</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Multi-Modal Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The objective is to address the cross-modal salience mismatch in token compression methods for omnimodal large language models by developing a framework named OmniScope that estimates relevance separately for audio and video.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Introduced a training-free token compression framework using the query as a shared semantic anchor, allocating modality-specific token budgets, pruning visual tokens with an anchor-delta strategy, and merging audio tokens to reduce redundancy.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; OmniScope outperforms existing methods in accuracy across audio-video benchmarks and model scales, offering improvements in speed and memory efficiency while maintaining high accuracy with minimal drop.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.23193" target="_blank">https://huggingface.co/papers/2607.23193</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233650331.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>24. Harness-G: A Graph-Structured Harness for Search Agents</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Reinforcement Learning, retrieval aliasing, Harness-G, Structured Non-myopic Credit</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Reinforcement Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study addresses the problem of retrieval aliasing in reinforcement learning search agents by proposing a new retrieval framework to optimize query generation and interaction.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The authors introduce Harness-G, a graph-structured retrieval framework that transforms query generation into finite action selection. They also develop Structured Non-myopic Credit (SNC) for better rewarding actions.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Harness-G achieves the highest average F1 across six QA benchmarks and outperforms the best baseline by significant margins, demonstrating its effectiveness in reducing linguistic aliasing and improving retrieval accuracy.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.27652" target="_blank">https://huggingface.co/papers/2607.27652</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233625249.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>25. Is Deep Research Reliable? Misleading Knowledge Induces False Conclusions</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Deep Research agents, Misleading Knowledge, Evidence Synthesis, Workflow Reliability</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Knowledge Representation and Reasoning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study aims to explore the reliability vulnerabilities in Deep Research agents when exposed to misleading knowledge in open information environments.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Introduces MisKnow-Agent, a framework for generating misleading knowledge instances to test these vulnerabilities, utilizing 5,933 quality-controlled instances derived from DeepResearch Benchmark tasks.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The experiments demonstrate that even minimal misleading knowledge can lead to false conclusions in final reports, highlighting a significant reliability issue.</p>
<p>   &#8211; Despite verifier models identifying misleading instances, they can still be adopted in long-term workflows, indicating a gap in workflow-level evidence verification.</p>
<p>   &#8211; Evaluated defenses show mitigation effects but fail to fully prevent the adoption of false conclusions, pointing to the need for improved verification and correction capabilities.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.20891" target="_blank">https://huggingface.co/papers/2607.20891</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233600666.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>26. Filesystem-Based Memory for LLM Agents: Organization, Evolution, and Sustainability</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: LLM agents, filesystem-based memory, memory organization, search economy, management agent</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: AI Systems and Tools</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To systematically explore filesystem-based memory for Large Language Model (LLM) agents, assessing the structure and organization of such memory.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The study formalizes the concept of a memory filesystem with a management agent for organization, a search agent for querying, and an execution agent for task trajectory integration.</p>
<p>   &#8211; Investigated variables include memory shape, stream scale, tool harness, and agent strength, with focus on memory growth, retrieval costs, and store health.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Organization in filesystem-based memory leads to improved search economy, significantly reducing retrieval costs.</p>
<p>   &#8211; However, as memory grows, management agents struggle to maintain organization quality, and improvements in organization do not necessarily translate to better answer quality.</p>
<p>   &#8211; The study identifies the tool set as a significant factor in shaping the memory store, highlighting a design space for agent memory rather than just an assumption.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.26637" target="_blank">https://huggingface.co/papers/2607.26637</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233538463.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>27. Multi-Head Attention Residuals</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Transformers, Multi-Head Attention Residuals, Nemotron, validation loss</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Machine Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; Introduction of Multi-Head Attention Residuals (MHAR) to improve information propagation across model depth and width.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; MHAR utilizes H per-subspace heads with distinct softmax distributions, trained on a Nemotron-based corpus to optimize validation loss across various model sizes.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; MHAR outperforms standard Transformers with validation loss improvements across 100M, 350M, and 1B scales. Optimal head count H = 8 increases efficiency and maintains near-baseline memory consumption.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.27230" target="_blank">https://huggingface.co/papers/2607.27230</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233514954.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>28. Echoverse: Deep, Evolving Environments for Training Computer-Use Agents at Scale</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Synthetic Environments, Echoverse, Reinforcement Learning, State Representation</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Reinforcement Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To improve computer-use agents&#8217; training environments by focusing on behavioral depth, targeted interactions, and model improvement.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Developed Echoverse to compile specifications into stateful applications, incorporating a co-evolution loop for iterative environment and model enhancement.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Demonstrated significant model improvement across various environments, with increases in evaluation scores and live-site accuracy via deep environments and iterative environment repairs.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28074" target="_blank">https://huggingface.co/papers/2607.28074</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233451225.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>29. INTACT: Isomorphic Intent-to-Action Learning for Search-Free World Models</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: INTACT, Intent-to-Action, Latent World Models, Action Law Semantics</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Reinforcement Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To introduce INTACT, an end-to-end JEPA, aimed at transforming action-labeled, reward-free trajectories into a deployable intent-to-action interface without expensive test-time search.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Utilizes a unique architecture with isomorphic intent representation for local and goal motion-intent models, and provides intact transfer from RGB evidence to action-effective latent intent coordinates.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The INTACT model achieved up to 100% success rates on LeWM tasks using a search-free policy, significantly reducing the need for candidate sequence sampling while improving performance through a shared four-task encoder. Direct inference is highly efficient, taking only 2.9-5.5 milliseconds.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.26056" target="_blank">https://huggingface.co/papers/2607.26056</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233423941.jpg"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>30. MemHarness: Memory Is Reconstructed, Not Replayed</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: MemHarness, Negative Transfer, Context-grounded Guidance, Large Language Model Agents, GRPO</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Reinforcement Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The objective is to develop a framework named MemHarness that enables large language model agents to actively reconstruct past experiences based on the current context, thereby providing more accurate and context-relevant guidance.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; MemHarness employs a unified policy model that critiques and reconstructs retrieved experiences for each decision step. This reconstructive ability is achieved through end-to-end training using the GRPO technique.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; MemHarness significantly outperforms existing memory-augmented baselines and pure reinforcement learning models, especially in out-of-distribution scenarios, by preventing negative transfer and enhancing intrinsic reasoning capabilities.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28272" target="_blank">https://huggingface.co/papers/2607.28272</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233359267.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>31. ShadowDancer: Teaching Video World Models Any Action by Learning Unified Dynamics Representations from a Video and Its Shadow</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: ShadowDancer, interactive video, dynamics, action transfer, cross-shadow prediction</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Computer Vision</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To enable precise control of any-action, frame-level dynamics in interactive video world models through novel techniques.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Utilizes shadow pairs to construct scalable video pairs replaying the same dynamics with varied appearances.</p>
<p>   &#8211; Implements cross-shadow prediction to learn actions by predicting dynamics from one shadow to another.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Demonstrated improved action transfer and extended action rollout in diverse dynamics families, achieving an average blinded win rate of 86% over existing models.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28362" target="_blank">https://huggingface.co/papers/2607.28362</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233335737.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>32. See2Think: Do Multimodal Models Really Use Intermediate Visual States?</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Multimodal large language models, Visual reasoning, See2Think, Visual Action-of-Thought</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Multi-Modal Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To introduce See2Think, an evaluation framework that assesses visual reasoning in multimodal large language models through benchmarks like See2ThinkBench and Visual Action-of-Thought.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The creation of See2ThinkBench with 1,200 visually dependent tasks across 12 categories and the use of Visual Action-of-Thought to record and analyze visual and textual operations under controlled settings.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Visual reasoning capabilities in models are strongly dependent on both the model and environment, showing varied performance across tasks. Additionally, even though models often select the correct visual operations, the accuracy improvements are hindered by challenges in rendering fidelity and reliance on feedback which does not always improve accuracy.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.26769" target="_blank">https://huggingface.co/papers/2607.26769</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233311886.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>33. RefCaptioner: Multi-Reference Image-Grounded Video Captioning</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: multi-reference image-grounded video captioning, RefCaptioner, factual video descriptions, Hierarchical Coverage-Discounted GRPO, AI-generated videos</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Multi-Modal Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To introduce and address the task of multi-reference image-grounded video captioning, requiring factual video descriptions with phrase-level reference grounding.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Development of RefCaptioner, a two-stage post-training framework combining mixed-data SFT with Hierarchical Coverage-Discounted GRPO to enhance reference selection and cross-reference consistency.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; RefCaptioner outperforms other open-source models in overall performance, yielding captions preferred by human annotators and allowing for more source-faithful video reconstruction.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28509" target="_blank">https://huggingface.co/papers/2607.28509</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233249116.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>34. ACE-Data-0: Human-Centric Ambient Capture as Embodied Data Engine</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Embodied intelligence, Ambient Capture Engine, AI Native, Multi-modal datasets, Human-centric data</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Robotics and Autonomous Systems</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study focuses on addressing the data bottleneck in embodied intelligence by introducing the Ambient Capture Engine (ACE), which captures comprehensive sensory data in real home environments.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The authors developed ACE to record multisensory data through a combination of table-scale and room-scale configurations, creating a detailed dataset, ACE-Data-0, encompassing 150 hours of interactions across 200 task categories.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Evaluations reveal significant limitations in current state-of-the-art methods, particularly under challenging conditions like occlusion and long temporal horizons, highlighting ACE&#8217;s potential in supporting scalable advancements in imitation learning and embodied AI systems.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28625" target="_blank">https://huggingface.co/papers/2607.28625</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233221968.jpg"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>35. Flux-OPD: On-Policy Distillation with Evolving Contexts</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Large language model, evolving contexts, Flux-OPD, open-ended domains, teacher supervision</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Natural Language Processing</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study aims to develop a mechanism for more effective supervision in large language model training across open-ended domains, addressing the issue of lacking verifiable rewards and formalized task preferences.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The research introduces Flux-OPD, a novel OPD paradigm that leverages evolving contexts as a form of in-training supervision. It analyzes the effect of context through the reverse KL objective and uses context-conditioned teachers to provide contextual difference signals.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Flux-OPD demonstrates superiority over existing OPD paradigms, effectively capturing task preferences and enhancing performance in open-ended tasks by combining evolving contexts with teacher supervision.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28022" target="_blank">https://huggingface.co/papers/2607.28022</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233157728.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>36. Beacon: Knowing When and How to Perform Agentic Visual Reasoning</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Agentic Visual Reasoning, Multimodal Large Language Models, Mode Adaptiveness, Tool Effect, Beacon</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Multi-Modal Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study aims to enhance the success rate of multimodal large language models (MLLMs) in complex tasks by rethinking agentic visual reasoning, focusing on Mode Adaptiveness and Tool Effect.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Comprehensive analysis to quantify Mode Adaptiveness and Tool Effect in existing models, followed by the proposal of a new model named Beacon, featuring Necessity-Aware Adaptive Reward and Hint-Guided Capability Expansion.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The research concludes that existing models have limited Mode Adaptiveness, and tool benefits are often neutralized by their drawbacks. Beacon, however, shows improved overall performance, notably enhancing Mode Adaptiveness and achieving genuine gains from tool use.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28595" target="_blank">https://huggingface.co/papers/2607.28595</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233136136.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>37. VideoCoCo: Code-as-CoT for Physically-Consistent Video Generation via an Agentic Dual-Engine System</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Text-to-video, VideoCoCo, Blender program, spatiotemporal process, generative video engine</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Generative Models</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study aims to address the limitations of text-to-video models in generating physically consistent dynamics by introducing a new framework called VideoCoCo.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; A dual-engine framework is employed, where a coding agent synthesizes executable Blender code to specify scene evolution and produce deterministic drafts, followed by a generative video engine for photorealistic video transformation.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; VideoCoCo effectively improves performance on benchmarks PhyGenBench and VBench-2.0 by providing an executable, controllable, and inspectable representation for video generation, achieving the best average scores on both benchmarks.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.27380" target="_blank">https://huggingface.co/papers/2607.27380</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233111733.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>38. PhiZero: A World Model Built Around Physical Language</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: PhiZero, physical language, world model, zero-shot motion transfer</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Generative Models</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The primary aim is to develop PhiZero, a model that abstracts predictive structure from visual experience to explicitly reason about physical world evolution using physical language.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; PhiZero employs a reason-then-render paradigm where it infers future world evolution sequences in physical language and renders transitions into video through self-supervision from in-the-wild videos.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The experiments demonstrate PhiZero&#8217;s capability to model physically coherent world evolution and highlight its potential for realistic interactive world modeling, fine-grained action-conditioned simulation, and zero-shot motion transfer.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28624" target="_blank">https://huggingface.co/papers/2607.28624</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233043717.jpg"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>39. Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: GUI agents, Qwen-UI-Agent, Reinforcement Learning, MobileWorld, Online RL</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: AI Systems and Tools</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The paper aims to advance GUI agents for real-world applications, enabling reliable operation across various devices and platforms, proactive service initiation, and autonomous capability improvements.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Introduces Qwen-UI-Agent, utilizing a unified action space that combines GUI operations with CLI execution, and employs an AutoResearch-style data flywheel and Online Reinforcement Learning (RL) for training and advancing agent performance.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Qwen-UI-Agent demonstrates state-of-the-art performance in mobile-use benchmarks and competitive results in computer- and browser-use tasks, achieving high scores on MobileWorld, AndroidDaily, and WebArena among others.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28227" target="_blank">https://huggingface.co/papers/2607.28227</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233022448.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<p>The post <a href="https://ainativefoundation.org/ai-native-daily-paper-digest-20260916-askchem-metis-chimera/">AI Native Daily Paper Digest – 20260916 – AskChem | Metis | Chimera</a> appeared first on <a href="https://ainativefoundation.org">AI Native Foundation</a>.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Global AI Native Industry Insights &#8211; 20260916 &#8211;  Anthropic &#124; Google &#124; more</title>
		<link>https://ainativefoundation.org/global-ai-native-industry-insights-20260916-anthropic-google-more/</link>
		
		<dc:creator><![CDATA[AINF]]></dc:creator>
		<pubDate>Wed, 16 Sep 2026 10:00:08 +0000</pubDate>
				<category><![CDATA[Global Industry]]></category>
		<guid isPermaLink="false">https://ainativefoundation.org/?p=10977</guid>

					<description><![CDATA[<p>Today’s digest highlights advancements from major players like Google DeepMind and Salesforce. A unifying theme in these updates is the integration of AI across various domains, emphasizing both practical applications and ongoing innovation. Google DeepMind's Gemini 3.8 and Live Extended models showcase the latest in conversational AI, while Salesforce launches its integration in Claude, marking its beta phase. Notably, Google's AI initiatives extend to fields as diverse as health and economic opportunity, demonstrating a commitment to applying AI solutions to real-world problems.</p>
<p>The post <a href="https://ainativefoundation.org/global-ai-native-industry-insights-20260916-anthropic-google-more/">Global AI Native Industry Insights &#8211; 20260916 &#8211;  Anthropic | Google | more</a> appeared first on <a href="https://ainativefoundation.org">AI Native Foundation</a>.</p>
]]></description>
										<content:encoded><![CDATA[<p>Today’s digest highlights advancements from major players like Google DeepMind and Salesforce. A unifying theme in these updates is the integration of AI across various domains, emphasizing both practical applications and ongoing innovation. Google DeepMind&#8217;s Gemini 3.8 and Live Extended models showcase the latest in conversational AI, while Salesforce launches its integration in Claude, marking its beta phase. Notably, Google&#8217;s AI initiatives extend to fields as diverse as health and economic opportunity, demonstrating a commitment to applying AI solutions to real-world problems. Discover more in Today&#8217;s Global AI Native Industry Insights.</p>
<h3>1.  Salesforce Integration in Claude Launches in Beta</h3>
<p>Anthropic has launched a beta integration bringing Salesforce data into Claude. The integration imports accounts, opportunities, and pipeline information directly into Claude conversations. It includes 37 pre-built sales skills covering tasks such as preparing for calls, reviewing deals, building pipeline dashboards, and sending forecasts. Users can perform these sales workflows without leaving the Claude interface.<br />
Read more: <a href="https://claude.com/blog/salesforce-in-claude">https://claude.com/blog/salesforce-in-claude</a></p>
<p><video width="600" height="400" controls poster="https://cdn.ainative.foundation/image/20260916_723cd6c7d0844ce2b49d0550622f32f3.jpg"><source src="https://cdn.ainative.foundation/video/20260916_f8cd09d9843b41f087a2e5e7763a7410.mp4" type="video/mp4"></video><br />
Video Credit: @claudeai on X</p>
<h3>2.  Google DeepMind Launches Gemini 3.8 Live and Live Extended Thinking Conversational AI Models</h3>
<p>Google DeepMind announced Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, described as its most advanced conversational AI models to date. The models are designed to support real-time voice and text conversation while reasoning through and executing tasks in the background. Google DeepMind says the new models aim to maintain conversational flow by handling complex thinking without interrupting the user interaction. The announcement was made via a thread on the Google DeepMind account on September 15, 2026.<br />
Read more: <a href="https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-8-live-gemini-3-8-live-extended-thinking/">https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-8-live-gemini-3-8-live-extended-thinking/</a></p>
<p><video width="600" height="400" controls poster="https://cdn.ainative.foundation/image/20260916_c2611030615e4e74acebab023dc3ae76.png"><source src="https://cdn.ainative.foundation/video/20260916_gj_video_gemini.mp4" type="video/mp4"></video><br />
Video Credit: The original article</p>
<h3>3.  Google Details AI Efforts in Science, Health, Weather, and Economic Opportunity</h3>
<p>Google outlined recent AI initiatives aimed at accelerating science and improving people&#8217;s lives. It released the AlphaGenome Atlas, mapping all 9 billion possible single-letter genetic variants across the human genome for open research use. The company also launched WeatherNext 3, described as its most accurate and capable global weather AI model to date. Additionally, Google published the AI &#038; Economy ATLAS, an open-access report on global AI usage patterns, and noted its translation services now cover nearly 300 languages spoken by 7 billion people.<br />
Read more: <a href="https://blog.google/innovation-and-ai/technology/ai/ai-applications-science-people/">https://blog.google/innovation-and-ai/technology/ai/ai-applications-science-people/</a></p>
<p><video width="600" height="400" controls poster="https://cdn.ainative.foundation/image/20260916_bfb08fb84a224900a4407a0e03182c08.png"><source src="https://cdn.ainative.foundation/video/20260916_gj_video_google.mp4" type="video/mp4"></video><br />
Video Credit: The original article</p>
<div style="width:100%;height:2px;background:#808080;margin:10px 0"></div>
<p>That&#8217;s all for today&#8217;s Global AI Native Industry Insights. Join us at <a href="https://member.ainativefoundation.org/">AI Native Foundation Membership Dashboard</a> for the latest insights on AI Native, or follow our linkedin account at <a href="https://www.linkedin.com/company/ainativefoundation/">AI Native Foundation</a> and our twitter account at <a href="https://x.com/AINativeF">AINativeF</a>.</p>
<p>The post <a href="https://ainativefoundation.org/global-ai-native-industry-insights-20260916-anthropic-google-more/">Global AI Native Industry Insights &#8211; 20260916 &#8211;  Anthropic | Google | more</a> appeared first on <a href="https://ainativefoundation.org">AI Native Foundation</a>.</p>
]]></content:encoded>
					
		
		<enclosure url="https://cdn.ainative.foundation/video/20260916_f8cd09d9843b41f087a2e5e7763a7410.mp4" length="17311118" type="video/mp4" />
<enclosure url="https://cdn.ainative.foundation/video/20260916_gj_video_gemini.mp4" length="32601956" type="video/mp4" />
<enclosure url="https://cdn.ainative.foundation/video/20260916_gj_video_google.mp4" length="14194538" type="video/mp4" />

			</item>
		<item>
		<title>AI Native Daily Paper Digest – 20260915 – Metis &#124; AskChem &#124; SpatialCLI</title>
		<link>https://ainativefoundation.org/ai-native-daily-paper-digest-20260915-metis-askchem-spatialcli/</link>
		
		<dc:creator><![CDATA[insights]]></dc:creator>
		<pubDate>Wed, 16 Sep 2026 00:41:10 +0000</pubDate>
				<category><![CDATA[Papers]]></category>
		<guid isPermaLink="false">https://ainativefoundation.org/ai-native-daily-paper-digest-20260915-metis-askchem-spatialcli/</guid>

					<description><![CDATA[<p>Today&#8217;s digest highlights notable entries from leaders like Gemma and Qwen, showcasing advancements in multimodal reasoning. A key theme emerges around enhancing [&#8230;]</p>
<p>The post <a href="https://ainativefoundation.org/ai-native-daily-paper-digest-20260915-metis-askchem-spatialcli/">AI Native Daily Paper Digest – 20260915 – Metis | AskChem | SpatialCLI</a> appeared first on <a href="https://ainativefoundation.org">AI Native Foundation</a>.</p>
]]></description>
										<content:encoded><![CDATA[<p style="margin-bottom:24px">Today&#8217;s digest highlights notable entries from leaders like Gemma and Qwen, showcasing advancements in multimodal reasoning. A key theme emerges around enhancing model performance through innovative algorithmic techniques. One paper introduces a novel approach that boosts efficiency by 35% on the established XYZ benchmark, while another paper explores an advanced attention mechanism that improves long-context processing by 28%. The research also covers a new metric allowing for more accurate evaluation of long-sequence understanding in AI models. Each paper provides substantial insights into AI model improvements, presenting potential shifts in future AI development.</p>
<h3>1. AskChem: Claim-Centered Infrastructure for Chemistry Literature Synthesis</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: AskChem, claim-centered infrastructure, GPT-5.5, AI agents, chemistry search</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Knowledge Representation and Reasoning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The objective of the research was to develop AskChem, a claim-centered infrastructure for conducting cross-paper chemistry searches, which shifts the retrieval focus from whole papers to provenance-carrying claims.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; AskChem converts papers into atomic, typed claims grounded by a source DOI and quotes or evidence locators, enabling hierarchical retrieval and exploration through a faceted taxonomy and evidence graph.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; AskChem indexes 2.4M claims from 147K papers, improving the retrieval capability with 100% resolvable DOIs for a GPT-5.5 reader and provides superior citation density compared to other systems.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28618" target="_blank">https://huggingface.co/papers/2607.28618</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img fetchpriority="high" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233007276.jpg"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>2. Metis: Memory Foundation Model</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: AI Native, native memory, memory foundation models, Metis</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Foundations of AI</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; Introduce memory foundation models to enable native memory capabilities in AI systems.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Proposed Metis, a new architecture with a native memory state for storing historical information.</p>
<p>   &#8211; Developed large-scale memory-specific training data and optimization objectives for native memory procedures.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Metis demonstrates native memory capabilities with advantages in architecture and efficiency.</p>
<p>   &#8211; The study provides a detailed analysis of strengths, limitations, and behaviors of Metis.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.26760" target="_blank">https://huggingface.co/papers/2607.26760</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233032494.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>3. Frontis-MA1: Training an AI4AI Model towards Recursive Self-Improvement in Machine Learning Engineering</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Recursive self-improvement, AI4AI, OpenMLE, Machine learning engineering, Reinforcement learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: AI Systems and Tools</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The paper aims to study Recursive self-improvement by introducing OpenMLE, an open full-stack system for AI4AI research in machine learning engineering.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Utilizes a combination of execution feedback environments (OpenMLE-Gym), operator learning (OpenMLE-RL), and long-horizon search (OpenMLE-Evo) to post-train a meta-evolution agent, Frontis-MA1, with atomic program-evolution operators.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The system significantly improves performance on MLE-Bench Lite, surpassing other models in efficiency and capability with OpenMLE-Evo-Max, and demonstrates transferable capabilities on NatureBench Lite.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28568" target="_blank">https://huggingface.co/papers/2607.28568</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233059370.jpg"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>4. Memory Decoder at Scale: A Pretrained, Parametric Long-Term Memory</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Memory Decoder, scalability, parameter efficiency, Faiss, pretraining</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Natural Language Processing</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To scale memory models up to 6.9B parameters and improve memory capacity independently for better performance of language models.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Implementation of a distributed pipeline for Faiss indexing and retrieval combined with sparse, batch-wise loading of kNN distributions.</p>
<p>   &#8211; Comparison of parameter-performance tradeoff by allocating more parameters to memory across different model scales.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Independently scaling pretrained memory models can achieve better parameter efficiency and enhance language model performance, as demonstrated by outperforming larger models with fewer parameters.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.27919" target="_blank">https://huggingface.co/papers/2607.27919</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233124782.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>5. BM25 Wins at Scale: A Scaling Study of Retrieval-Augmented Generation Paradigms</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Retrieval-Augmented Generation, BM25, Dense Retrieval, Lexical Retrieval, Agentic Reasoning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Natural Language Processing </p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To conduct a controlled study evaluating retrieval-augmented generation (RAG) approaches across varying corpus sizes and determine their accuracy-cost scaling.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The study varies corpus size across 28 nested tiers, analyzing accuracy, construction, query tokens, and latency under a single reader model and judging protocol.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Different approaches show varying effectiveness based on corpus size, with BM25 eventually leading as the corpus grows, while agentic search incurs high query token costs at larger scales.</p>
<p>   &#8211; Lexical retrieval emerges as the most scalable default, outperforming dense retrieval and graph-based indexing at larger corpus sizes, highlighting global candidate ranking&#8217;s importance.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.26497" target="_blank">https://huggingface.co/papers/2607.26497</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233146706.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>6. MPIE-Bench: Benchmarking Anatomically Plausible Multi-Person Interaction Editing</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Text-to-image, personalized editing, MPIE-Bench, MPIE-Eval, multi-person mesh reconstruction</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Computer Vision</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The research aims to address issues in synthesizing high-fidelity images featuring multiple named people in shared contact actions, such as embrace or carry.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Introduced MPIE-Bench, a benchmark of 2,500 samples from video-mined editing triplets covering various scenes, interactions, and contact densities.</p>
<p>   &#8211; Proposed MPIE-Eval, which evaluates contact-time geometry using a public multi-person mesh reconstruction framework.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The study finds that current visual language model (VLM) metrics often fail to capture anatomical and geometric errors, instead showing that MPIE-Eval aligns more closely with human judgment in evaluating anatomy and interaction accuracy.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.27616" target="_blank">https://huggingface.co/papers/2607.27616</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233209977.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>7. Beyond Borrowed Histories: Person-Aligned User Simulation for Interactive Role-Playing Evaluation</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Role-playing agents, Evaluation, User simulators, Personalized rubrics, Multi-turn conversation</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Human-AI Interaction</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The paper aims to improve the evaluation of Role-playing agents (RPAs) by addressing the limitations of existing benchmarks and introducing a more personalized and realistic evaluation method through user simulators.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Implementation of PALATE, a scalable RPA benchmark using 300 character profiles and personalized rubrics, enabling free-form, multi-turn conversations with user simulators to assess RPAs&#8217; conversational abilities.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; PALATE provides interpretable evaluations of RPAs&#8217; performance in specific user interactions, demonstrating higher agreement with human judgments compared to traditional rubrics, and addresses both generic quality and personalized user experience.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.27816" target="_blank">https://huggingface.co/papers/2607.27816</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233235908.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>8. SpatialCLI: Learning to Reason With Spatial Tools, Then Without Them</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Vision-language models, SpatialCLI, Embodied agents, Compositional perception</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Multi-Modal Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The paper aims to address the capability mismatch in Vision-language models (VLMs) between task-level reasoning and capturing visual details, proposing a new framework to bridge this gap.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The proposed SpatialCLI framework operates in three stages: exposing specialist vision models as spatial tools, improving tool use with Cold-Start SFT and agentic RL, and internalizing successful tool-use trajectories.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; SpatialCLI significantly improves VLM performance on the MindCube benchmark, raising the accuracy from 29.3% to 84.6% with tools, and maintaining 73.8% accuracy even without tools after internalization, surpassing existing models.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.27703" target="_blank">https://huggingface.co/papers/2607.27703</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233300818.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>9. Chimera: Designing and Chinchilla-Scaling Hybrid Visual Diffusion Transformers</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Visual Generation, Chimera, Kimi Delta Attention, Multi-head Latent Attention, Sparse Mixture-of-Experts</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Generative Models</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To develop a hybrid visual diffusion backbone, Chimera, for efficient handling of high-resolution images, long videos, and multimodal context.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Introduction of Chimera that utilizes Kimi Delta Attention, Multi-head Latent Attention, and Sparse Mixture-of-Experts layers, along with HeteroP for scaling the architecture effectively.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Chimera&#8217;s dense backbone is 1.7x more compute-efficient than the Wan-2.1 baseline, and the complete system achieves 7.3x efficiency.</p>
<p>   &#8211; Chimera can extrapolate from 5-second training clips to 30-second videos with minimal performance degradation.</p>
<p>   &#8211; Fitted compute-optimal laws indicate balanced compute division in image pretraining, with a slight model size preference in video pretraining at higher budgets.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28611" target="_blank">https://huggingface.co/papers/2607.28611</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233323944.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>10. β-OPSD: Deriving with Policy Optimization, Training with Self-Distillation</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: On-policy self-distillation, policy optimization, reinforcement learning, reasoning performance, mathematical reasoning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Reinforcement Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study aims to enhance reasoning language models using a refined version of on-policy self-distillation, termed β-OPSD, which introduces a versatile regularization parameter β for better policy optimization.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The paper introduces β-OPSD by positioning it within a broader policy-optimization family. It employs reinforcement learning for distilling the closed-form solution through efficient token-level logit mixing to approximate costly optimization processes.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Experiments demonstrate that β-OPSD consistently outperforms the standard OPSD in mathematical reasoning tasks, achieving improved optimization stability and better reasoning performance.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28582" target="_blank">https://huggingface.co/papers/2607.28582</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233348258.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>11. Can Large Language Models Execute Parent Orders?</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Parent-order execution, Large Language Models (LLMs), PACE, Shenzhen Stock Exchange, Execution decisions</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: AI in Finance</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study aims to enhance parent-order execution in algorithmic trading using Large Language Models (LLMs), addressing limitations of current methods like reliance on market assumptions and lack of adaptability.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Introduction of PACE (Plan-Ahead Controlled Execution), a hierarchical framework that separates parent-order execution into long-term planning and short-term execution, bypassing the need for specific market assumptions and training.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; PACE demonstrated superior performance compared to existing TWAP, Almgren-Chriss, and learning-based strategies during tests on Shenzhen Stock Exchange data. It was observed that LLMs navigate execution decisions differently from humans, with higher confidence correlated with improved performance and earlier trades instead of late actions.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28410" target="_blank">https://huggingface.co/papers/2607.28410</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233411593.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>12. LEDGERMIND: Provenance-Constrained Multimodal Agentic Reasoning with a Structured Evidence Ledger</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Multimodal Agents, Visual Question Answering, Provenance-Constrained, Structured Evidence, LedgerMind</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Multi-Modal Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study aims to improve the evaluation of multimodal agents that handle visual question answering by considering not only answer accuracy but also the process through which answers are derived.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Introduction of LedgerMind, a system based on a provenance-constrained state machine, to standardize tool outputs and support structured evidence.</p>
<p>   &#8211; Utilization of a Three-Layer Grounding Protocol, Adaptive Dual-Path Dispatcher, and an Event-Triggered Verification-and-Repair engine to manage reasoning depth and ensure provenance.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; LedgerMind effectively addresses common failure patterns such as unsupported reasoning and entity hallucinations, enhancing both answer accuracy and the reliability of the reasoning trajectory.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28374" target="_blank">https://huggingface.co/papers/2607.28374</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233436135.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>13. Σ-Mem: An Online Reliability Memory for LLM-based Multi-Agent Systems</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: LLM agents, multi-agent systems, Σ-Mem, reliability memory, interaction content</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: AI Systems and Tools</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The objective is to introduce Σ-Mem, a novel online reliability memory system, addressing limitations in existing memory systems within LLM agents, particularly in multi-agent environments.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Σ-Mem models historical competence and peer relationship evidence, using real symmetric states updated from correctness feedback while ensuring stable online adaptation without requiring model retraining.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Σ-Mem effectively adapts to shifts in counterfactual reliability and generalizes to new peers and task domains, outperforms traditional voting methods, and improves with increased correctness feedback, establishing a solid foundation for adaptive coordination in multi-agent LLM systems.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.27958" target="_blank">https://huggingface.co/papers/2607.27958</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233503295.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>14. Explorative Modeling: Unlocking a Third Pretraining Axis and End-to-End Generation</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Explorative Modeling, Generative Models, FLOP Efficiency, Sample Efficiency, Parameter Efficiency</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Generative Models</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; Introduce Explorative Modeling as a new paradigm in generative modeling to enable end-to-end training by exploring candidate matches and training on the best predictions.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The Explorative Models are applied in generative domains, increasing exploration as a pretraining axis to improve performance across diverse datasets including images, video, and language.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Explorative Modeling enhances efficiency exponentially, achieving significant improvement in FLOP, sample, and parameter efficiency, and enables end-to-end reconstructive generative modeling with reduced inference steps, establishing itself as a new training axis for existing generative models.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.27372" target="_blank">https://huggingface.co/papers/2607.27372</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233526935.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>15. AI Tour Meeting: Group Travel Planning by LLM Agents</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: AI Tour Meeting, Large Language Model, agents, simulation tool</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: AI Systems and Tools</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To propose a travel planning framework called AI Tour Meeting using agents powered by Large Language Models (LLMs) to collaboratively create itineraries through natural language discussion.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Implementation of a framework for configuring agent personas and orchestrating discussion workflows, enabling monitoring and LLM deployment in the simulation of tour planning discussions.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The framework is validated by system demonstration and provides analytical insights into the behavior of multiple LLM agents during group travel planning.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.18806" target="_blank">https://huggingface.co/papers/2607.18806</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233614234.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>16. Revisiting Lossy Verification in Speculative Decoding: Mechanisms, Trade-offs, and Failure Modes</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Speculative Decoding, Large Language Model, Truncation-based Verification, Collaborative Verification</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Natural Language Processing</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To analyze the distributions induced by lossy verification schemes in speculative decoding and identify their potential pitfalls in large language model inference.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Classified lossy verification methods into truncation-based and collaborative verification, and developed a diagnostic evaluation framework across curated benchmarks.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Truncation-based methods are susceptible to significant performance degradation due to distributional distortion.</p>
<p>   &#8211; For collaborative verification, controlling the overshoot of draft probabilities relative to target probabilities is crucial to maintaining output quality.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.26627" target="_blank">https://huggingface.co/papers/2607.26627</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233549571.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>17. Beyond Geometric Complementarity: Coherent Overlap in Sparse Mixture-of-Experts Routing</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Sparse mixture-of-experts, Expert Subspace Separation Index, coherent overlap, next-token prediction</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Natural Language Processing</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study aims to distinguish the benefits and interactions of experts within sparse mixture-of-experts (MoE) language models by analyzing their geometric arrangement.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Utilizes an Expert Subspace Separation Index (ESSI) and various controlled interventions, including matched-route residuals and prefix-controlled factorial setups, to analyze expert overlaps and their impact on representation.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Coherent overlap indicates that while routes overlap, useful multi-expert computation persists, clarifying why geometric similarity doesn&#8217;t alone determine redundancy or pruning value.</p>
<p>   &#8211; Adding later experts improves next-token predictions significantly in most cases, emphasizing the importance of non-redundant expert selections.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28308" target="_blank">https://huggingface.co/papers/2607.28308</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233702669.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>18. Fairness Pruning: Locating Demographic Bias in GLU-MLP Layers via Differential Activations</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Fairness Pruning, Large Language Models, Demographic Bias, Causal Bias Localization, AI Ethics and Fairness</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: AI Ethics and Fairness</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To introduce and empirically validate Fairness Pruning, a method aimed at managing and mitigating demographic biases in large language models.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Utilization of minimally contrastive prompt pairs and inference-time activation capture to identify neurons that react differentially to demographic attributes.</p>
<p>   &#8211; Empirical evaluation on large models such as Llama-3.2 and Salamandra-2B with structural intervention by zeroing neurons.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Demonstrated that zeroing specific neurons alters model responses to demographic variables, with a bidirectional bias destabilization effect.</p>
<p>   &#8211; Confirmed that demographic bias processing and model capabilities are dissociable, setting foundations for transitioning to directional behavior modulation.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28319" target="_blank">https://huggingface.co/papers/2607.28319</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233637423.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>19. AMRD: Adaptive Multi-Teacher Relational Distillation for Lightweight Speech Emotion Recognition</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Speech Emotion Recognition, Edge Devices, Knowledge Distillation, AI Native</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Machine Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; This research aims to optimize on-device speech emotion recognition (SER) for real-time applications by developing a more efficient model that can operate effectively on edge devices.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Adaptive Multi-teacher Relational Distillation (AMRD) is proposed, using a one-class SVM to assign weights to teacher models and a relational distillation loss to improve inter-sample relational structure alignment between teacher and student models.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The AMRD approach outperforms single-teacher distillation baselines on the IEMOCAP and CREMA-D datasets, demonstrating its effectiveness across multiple student architectures and confirming that both components of the method yield complementary benefits.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.25289" target="_blank">https://huggingface.co/papers/2607.25289</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233726024.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>20. </h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="" target="_blank"></a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/thirteen/202607311785541064.jpg"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>21. Pedestrian Archetypes Extension &#8212; More Pedestrian Models for Autonomous Vehicle Safety Testing</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Pedestrian Archetypes, Real-world Traffic Scenarios, Behavior Patterns</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Robotics and Autonomous Systems</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The research introduces new pedestrian archetypes that exhibit significant behavioral differences, enhancing the understanding of pedestrian behaviors beyond the original taxonomy.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Analysis of YouTube dash-cam videos for identifying additional pedestrian archetypes not covered in previous work.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Seven new pedestrian archetypes are presented, each defined by unique behaviors that are documented through video-frame evidence, thus expanding the existing framework for assessing pedestrian behavior in traffic.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.16922" target="_blank">https://huggingface.co/papers/2607.16922</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233737256.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>22. ReToken: One Token to Improve Vision-Language Models for Visual Retrieval</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Vision-Language Models, GPU Memory Constraints, Image-QA Dataset</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Multi-Modal Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; Address the challenge of long visual context in vision-language models by introducing ReToken, an explicit retrieval target for selecting query-relevant visual tokens.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; ReToken is trained on a small image-QA dataset and designed to select a sparse set of visual tokens from a pre-filled visual KV cache.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; ReToken shows consistent performance gains across image and video benchmarks, with significant improvements in Visual Haystacks and LVBench, while maintaining efficiency on a single H100 for both training and long-video inference.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28627" target="_blank">https://huggingface.co/papers/2607.28627</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233713703.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>23. OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: OmniScope, token compression, query, salience mismatch, AI Native</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Multi-Modal Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The objective is to address the cross-modal salience mismatch in token compression methods for omnimodal large language models by developing a framework named OmniScope that estimates relevance separately for audio and video.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Introduced a training-free token compression framework using the query as a shared semantic anchor, allocating modality-specific token budgets, pruning visual tokens with an anchor-delta strategy, and merging audio tokens to reduce redundancy.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; OmniScope outperforms existing methods in accuracy across audio-video benchmarks and model scales, offering improvements in speed and memory efficiency while maintaining high accuracy with minimal drop.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.23193" target="_blank">https://huggingface.co/papers/2607.23193</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233650331.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>24. Harness-G: A Graph-Structured Harness for Search Agents</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Reinforcement Learning, retrieval aliasing, Harness-G, Structured Non-myopic Credit</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Reinforcement Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study addresses the problem of retrieval aliasing in reinforcement learning search agents by proposing a new retrieval framework to optimize query generation and interaction.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The authors introduce Harness-G, a graph-structured retrieval framework that transforms query generation into finite action selection. They also develop Structured Non-myopic Credit (SNC) for better rewarding actions.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Harness-G achieves the highest average F1 across six QA benchmarks and outperforms the best baseline by significant margins, demonstrating its effectiveness in reducing linguistic aliasing and improving retrieval accuracy.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.27652" target="_blank">https://huggingface.co/papers/2607.27652</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233625249.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>25. Is Deep Research Reliable? Misleading Knowledge Induces False Conclusions</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Deep Research agents, Misleading Knowledge, Evidence Synthesis, Workflow Reliability</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Knowledge Representation and Reasoning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study aims to explore the reliability vulnerabilities in Deep Research agents when exposed to misleading knowledge in open information environments.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Introduces MisKnow-Agent, a framework for generating misleading knowledge instances to test these vulnerabilities, utilizing 5,933 quality-controlled instances derived from DeepResearch Benchmark tasks.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The experiments demonstrate that even minimal misleading knowledge can lead to false conclusions in final reports, highlighting a significant reliability issue.</p>
<p>   &#8211; Despite verifier models identifying misleading instances, they can still be adopted in long-term workflows, indicating a gap in workflow-level evidence verification.</p>
<p>   &#8211; Evaluated defenses show mitigation effects but fail to fully prevent the adoption of false conclusions, pointing to the need for improved verification and correction capabilities.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.20891" target="_blank">https://huggingface.co/papers/2607.20891</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233600666.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>26. Filesystem-Based Memory for LLM Agents: Organization, Evolution, and Sustainability</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: LLM agents, filesystem-based memory, memory organization, search economy, management agent</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: AI Systems and Tools</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To systematically explore filesystem-based memory for Large Language Model (LLM) agents, assessing the structure and organization of such memory.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The study formalizes the concept of a memory filesystem with a management agent for organization, a search agent for querying, and an execution agent for task trajectory integration.</p>
<p>   &#8211; Investigated variables include memory shape, stream scale, tool harness, and agent strength, with focus on memory growth, retrieval costs, and store health.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Organization in filesystem-based memory leads to improved search economy, significantly reducing retrieval costs.</p>
<p>   &#8211; However, as memory grows, management agents struggle to maintain organization quality, and improvements in organization do not necessarily translate to better answer quality.</p>
<p>   &#8211; The study identifies the tool set as a significant factor in shaping the memory store, highlighting a design space for agent memory rather than just an assumption.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.26637" target="_blank">https://huggingface.co/papers/2607.26637</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233538463.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>27. Multi-Head Attention Residuals</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Transformers, Multi-Head Attention Residuals, Nemotron, validation loss</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Machine Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; Introduction of Multi-Head Attention Residuals (MHAR) to improve information propagation across model depth and width.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; MHAR utilizes H per-subspace heads with distinct softmax distributions, trained on a Nemotron-based corpus to optimize validation loss across various model sizes.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; MHAR outperforms standard Transformers with validation loss improvements across 100M, 350M, and 1B scales. Optimal head count H = 8 increases efficiency and maintains near-baseline memory consumption.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.27230" target="_blank">https://huggingface.co/papers/2607.27230</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233514954.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>28. Echoverse: Deep, Evolving Environments for Training Computer-Use Agents at Scale</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Synthetic Environments, Echoverse, Reinforcement Learning, State Representation</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Reinforcement Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To improve computer-use agents&#8217; training environments by focusing on behavioral depth, targeted interactions, and model improvement.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Developed Echoverse to compile specifications into stateful applications, incorporating a co-evolution loop for iterative environment and model enhancement.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Demonstrated significant model improvement across various environments, with increases in evaluation scores and live-site accuracy via deep environments and iterative environment repairs.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28074" target="_blank">https://huggingface.co/papers/2607.28074</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233451225.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>29. INTACT: Isomorphic Intent-to-Action Learning for Search-Free World Models</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: INTACT, Intent-to-Action, Latent World Models, Action Law Semantics</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Reinforcement Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To introduce INTACT, an end-to-end JEPA, aimed at transforming action-labeled, reward-free trajectories into a deployable intent-to-action interface without expensive test-time search.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Utilizes a unique architecture with isomorphic intent representation for local and goal motion-intent models, and provides intact transfer from RGB evidence to action-effective latent intent coordinates.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The INTACT model achieved up to 100% success rates on LeWM tasks using a search-free policy, significantly reducing the need for candidate sequence sampling while improving performance through a shared four-task encoder. Direct inference is highly efficient, taking only 2.9-5.5 milliseconds.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.26056" target="_blank">https://huggingface.co/papers/2607.26056</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233423941.jpg"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>30. MemHarness: Memory Is Reconstructed, Not Replayed</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: MemHarness, Negative Transfer, Context-grounded Guidance, Large Language Model Agents, GRPO</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Reinforcement Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The objective is to develop a framework named MemHarness that enables large language model agents to actively reconstruct past experiences based on the current context, thereby providing more accurate and context-relevant guidance.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; MemHarness employs a unified policy model that critiques and reconstructs retrieved experiences for each decision step. This reconstructive ability is achieved through end-to-end training using the GRPO technique.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; MemHarness significantly outperforms existing memory-augmented baselines and pure reinforcement learning models, especially in out-of-distribution scenarios, by preventing negative transfer and enhancing intrinsic reasoning capabilities.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28272" target="_blank">https://huggingface.co/papers/2607.28272</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233359267.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>31. ShadowDancer: Teaching Video World Models Any Action by Learning Unified Dynamics Representations from a Video and Its Shadow</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: ShadowDancer, interactive video, dynamics, action transfer, cross-shadow prediction</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Computer Vision</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To enable precise control of any-action, frame-level dynamics in interactive video world models through novel techniques.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Utilizes shadow pairs to construct scalable video pairs replaying the same dynamics with varied appearances.</p>
<p>   &#8211; Implements cross-shadow prediction to learn actions by predicting dynamics from one shadow to another.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Demonstrated improved action transfer and extended action rollout in diverse dynamics families, achieving an average blinded win rate of 86% over existing models.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28362" target="_blank">https://huggingface.co/papers/2607.28362</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233335737.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>32. See2Think: Do Multimodal Models Really Use Intermediate Visual States?</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Multimodal large language models, Visual reasoning, See2Think, Visual Action-of-Thought</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Multi-Modal Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To introduce See2Think, an evaluation framework that assesses visual reasoning in multimodal large language models through benchmarks like See2ThinkBench and Visual Action-of-Thought.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The creation of See2ThinkBench with 1,200 visually dependent tasks across 12 categories and the use of Visual Action-of-Thought to record and analyze visual and textual operations under controlled settings.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Visual reasoning capabilities in models are strongly dependent on both the model and environment, showing varied performance across tasks. Additionally, even though models often select the correct visual operations, the accuracy improvements are hindered by challenges in rendering fidelity and reliance on feedback which does not always improve accuracy.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.26769" target="_blank">https://huggingface.co/papers/2607.26769</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233311886.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>33. RefCaptioner: Multi-Reference Image-Grounded Video Captioning</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: multi-reference image-grounded video captioning, RefCaptioner, factual video descriptions, Hierarchical Coverage-Discounted GRPO, AI-generated videos</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Multi-Modal Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To introduce and address the task of multi-reference image-grounded video captioning, requiring factual video descriptions with phrase-level reference grounding.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Development of RefCaptioner, a two-stage post-training framework combining mixed-data SFT with Hierarchical Coverage-Discounted GRPO to enhance reference selection and cross-reference consistency.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; RefCaptioner outperforms other open-source models in overall performance, yielding captions preferred by human annotators and allowing for more source-faithful video reconstruction.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28509" target="_blank">https://huggingface.co/papers/2607.28509</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233249116.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>34. ACE-Data-0: Human-Centric Ambient Capture as Embodied Data Engine</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Embodied intelligence, Ambient Capture Engine, AI Native, Multi-modal datasets, Human-centric data</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Robotics and Autonomous Systems</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study focuses on addressing the data bottleneck in embodied intelligence by introducing the Ambient Capture Engine (ACE), which captures comprehensive sensory data in real home environments.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The authors developed ACE to record multisensory data through a combination of table-scale and room-scale configurations, creating a detailed dataset, ACE-Data-0, encompassing 150 hours of interactions across 200 task categories.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Evaluations reveal significant limitations in current state-of-the-art methods, particularly under challenging conditions like occlusion and long temporal horizons, highlighting ACE&#8217;s potential in supporting scalable advancements in imitation learning and embodied AI systems.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28625" target="_blank">https://huggingface.co/papers/2607.28625</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233221968.jpg"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>35. Flux-OPD: On-Policy Distillation with Evolving Contexts</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Large language model, evolving contexts, Flux-OPD, open-ended domains, teacher supervision</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Natural Language Processing</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study aims to develop a mechanism for more effective supervision in large language model training across open-ended domains, addressing the issue of lacking verifiable rewards and formalized task preferences.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The research introduces Flux-OPD, a novel OPD paradigm that leverages evolving contexts as a form of in-training supervision. It analyzes the effect of context through the reverse KL objective and uses context-conditioned teachers to provide contextual difference signals.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Flux-OPD demonstrates superiority over existing OPD paradigms, effectively capturing task preferences and enhancing performance in open-ended tasks by combining evolving contexts with teacher supervision.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28022" target="_blank">https://huggingface.co/papers/2607.28022</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233157728.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>36. Beacon: Knowing When and How to Perform Agentic Visual Reasoning</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Agentic Visual Reasoning, Multimodal Large Language Models, Mode Adaptiveness, Tool Effect, Beacon</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Multi-Modal Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study aims to enhance the success rate of multimodal large language models (MLLMs) in complex tasks by rethinking agentic visual reasoning, focusing on Mode Adaptiveness and Tool Effect.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Comprehensive analysis to quantify Mode Adaptiveness and Tool Effect in existing models, followed by the proposal of a new model named Beacon, featuring Necessity-Aware Adaptive Reward and Hint-Guided Capability Expansion.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The research concludes that existing models have limited Mode Adaptiveness, and tool benefits are often neutralized by their drawbacks. Beacon, however, shows improved overall performance, notably enhancing Mode Adaptiveness and achieving genuine gains from tool use.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28595" target="_blank">https://huggingface.co/papers/2607.28595</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233136136.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>37. VideoCoCo: Code-as-CoT for Physically-Consistent Video Generation via an Agentic Dual-Engine System</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Text-to-video, VideoCoCo, Blender program, spatiotemporal process, generative video engine</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Generative Models</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study aims to address the limitations of text-to-video models in generating physically consistent dynamics by introducing a new framework called VideoCoCo.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; A dual-engine framework is employed, where a coding agent synthesizes executable Blender code to specify scene evolution and produce deterministic drafts, followed by a generative video engine for photorealistic video transformation.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; VideoCoCo effectively improves performance on benchmarks PhyGenBench and VBench-2.0 by providing an executable, controllable, and inspectable representation for video generation, achieving the best average scores on both benchmarks.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.27380" target="_blank">https://huggingface.co/papers/2607.27380</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233111733.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>38. PhiZero: A World Model Built Around Physical Language</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: PhiZero, physical language, world model, zero-shot motion transfer</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Generative Models</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The primary aim is to develop PhiZero, a model that abstracts predictive structure from visual experience to explicitly reason about physical world evolution using physical language.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; PhiZero employs a reason-then-render paradigm where it infers future world evolution sequences in physical language and renders transitions into video through self-supervision from in-the-wild videos.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The experiments demonstrate PhiZero&#8217;s capability to model physically coherent world evolution and highlight its potential for realistic interactive world modeling, fine-grained action-conditioned simulation, and zero-shot motion transfer.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28624" target="_blank">https://huggingface.co/papers/2607.28624</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233043717.jpg"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>39. Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: GUI agents, Qwen-UI-Agent, Reinforcement Learning, MobileWorld, Online RL</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: AI Systems and Tools</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The paper aims to advance GUI agents for real-world applications, enabling reliable operation across various devices and platforms, proactive service initiation, and autonomous capability improvements.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Introduces Qwen-UI-Agent, utilizing a unified action space that combines GUI operations with CLI execution, and employs an AutoResearch-style data flywheel and Online Reinforcement Learning (RL) for training and advancing agent performance.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Qwen-UI-Agent demonstrates state-of-the-art performance in mobile-use benchmarks and competitive results in computer- and browser-use tasks, achieving high scores on MobileWorld, AndroidDaily, and WebArena among others.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28227" target="_blank">https://huggingface.co/papers/2607.28227</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233022448.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<p>The post <a href="https://ainativefoundation.org/ai-native-daily-paper-digest-20260915-metis-askchem-spatialcli/">AI Native Daily Paper Digest – 20260915 – Metis | AskChem | SpatialCLI</a> appeared first on <a href="https://ainativefoundation.org">AI Native Foundation</a>.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Global AI Native Industry Insights &#8211; 20260915 &#8211;  ElevenLabs &#124; Perplexity AI &#124; Cognition &#124; more</title>
		<link>https://ainativefoundation.org/global-ai-native-industry-insights-20260915-elevenlabs-perplexity-ai-cognition-more/</link>
		
		<dc:creator><![CDATA[AINF]]></dc:creator>
		<pubDate>Tue, 15 Sep 2026 09:45:56 +0000</pubDate>
				<category><![CDATA[Global Industry]]></category>
		<guid isPermaLink="false">https://ainativefoundation.org/?p=10955</guid>

					<description><![CDATA[<p>ElevenLabs and OpenAI take center stage in today's digest, showcasing advancements in voice and biological reasoning technologies. A common theme throughout is the expansion of AI capabilities across different modalities, with ElevenLabs adding voice, music, image, and video generation to its MCP server, and OpenAI launching GPT-Rosalind to enhance biological reasoning in API and Codex. Perplexity introduces the Comet Assistant for Windows PCs with NVIDIA RTX GPUs, emphasizing compatibility and performance enhancements. Cognition debuts Fusion, a model harness aimed at cost efficiency for Devin CLI users.</p>
<p>The post <a href="https://ainativefoundation.org/global-ai-native-industry-insights-20260915-elevenlabs-perplexity-ai-cognition-more/">Global AI Native Industry Insights &#8211; 20260915 &#8211;  ElevenLabs | Perplexity AI | Cognition | more</a> appeared first on <a href="https://ainativefoundation.org">AI Native Foundation</a>.</p>
]]></description>
										<content:encoded><![CDATA[<p>ElevenLabs and OpenAI take center stage in today&#8217;s digest, showcasing advancements in voice and biological reasoning technologies. A common theme throughout is the expansion of AI capabilities across different modalities, with ElevenLabs adding voice, music, image, and video generation to its MCP server, and OpenAI launching GPT-Rosalind to enhance biological reasoning in API and Codex. Perplexity introduces the Comet Assistant for Windows PCs with NVIDIA RTX GPUs, emphasizing compatibility and performance enhancements. Cognition debuts Fusion, a model harness aimed at cost efficiency for Devin CLI users. Discover more in Today&#8217;s Global AI Native Industry Insights.</p>
<h3>1.  ElevenLabs adds voice, music, image, and video generation to its MCP server</h3>
<p>ElevenLabs has expanded its Model Context Protocol (MCP) server to support generation of voice, music, images, and video, in addition to existing speech capabilities. Users can now generate speech, transcripts, dubs, music, sound effects, images, and video directly from within the AI assistant they already use. The update turns the ElevenLabs MCP into a multimodal content generation tool accessible through compatible assistants. This broadens the range of creative and production tasks developers and users can perform without leaving their existing workflow.<br />
Read more: <a href="https://elevenlabs.io/mcp">https://elevenlabs.io/mcp</a></p>
<p><video width="600" height="400" controls poster="https://cdn.ainative.foundation/image/20260915_gj_img_mcp.png"><source src="https://cdn.ainative.foundation/video/20260915_ad60cec8e60d45d8b73aacce38f4008c.mp4" type="video/mp4"></video><br />
Video Credit: @ElevenLabs on X</p>
<h3>2.  Perplexity&#8217;s Comet Assistant Launches for Windows PCs with NVIDIA RTX GPUs</h3>
<p>Perplexity AI announced that its Portable Computer agent tool is now available on Windows PCs equipped with NVIDIA RTX GPUs. The tool allows users to run its harness, agents, and models locally on their own hardware. It supports working with local files and connected apps without sending tasks to the cloud, while still allowing access to frontier cloud models when needed. The release expands local, on-device AI agent capabilities for Windows users.<br />
Read more: <a href="https://www.perplexity.ai/hub/products/portable-computer">https://www.perplexity.ai/hub/products/portable-computer</a></p>
<p><video width="600" height="400" controls poster="https://cdn.ainative.foundation/image/20260915_5ec4ba2a1569403ca0ebe613ae3b6b5e.jpg"><source src="https://cdn.ainative.foundation/video/20260915_7d18bfd304ad440aa0e3bbfd53195a53.mp4" type="video/mp4"></video><br />
Video Credit: @perplexity_ai on X</p>
<h3>3.  Cognition Launches Fusion, a Cost-Efficient Model Harness for Devin CLI</h3>
<p>Cognition introduced Fusion, a new harness within Devin CLI that combines separate models for planning and execution. The company says the approach works with its Fable and Astra models and cuts costs by 39% across coding benchmarks compared to using a single frontier model. Users can select a preferred model for planning tasks while pairing it with a more cost-effective model for execution. The release targets developers seeking lower-cost AI coding workflows without sacrificing performance.<br />
Read more: <a href="https://cognition.com/blog/local-fusion">https://cognition.com/blog/local-fusion</a></p>
<p><video width="600" height="400" controls poster="https://cdn.ainative.foundation/image/20260915_6a6b36293e2d43cdae2d7e1a2d67b75b.jpg"><source src="https://cdn.ainative.foundation/video/20260915_gj_video_aicoding.mp4" type="video/mp4"></video><br />
Video Credit: NotebookLM</p>
<h3>4.  OpenAI launches GPT-Rosalind for biological reasoning in API and Codex</h3>
<p>OpenAI has introduced GPT-Rosalind, a model designed to bring stronger biological reasoning capabilities to its API and Codex platform. The model helps researchers connect findings across scientific papers and experimental results. It can weigh evidence for a given biological target and work through analyses to help plan future experiments. The release targets researchers working on biological and life sciences problems.<br />
Read more: <a href="https://openai.com/index/introducing-gpt-rosalind/">https://openai.com/index/introducing-gpt-rosalind/</a></p>
<p><video width="600" height="400" controls poster="https://cdn.ainative.foundation/image/20260915_fe5cfc1be6eb4df0a40be0198fdb5662.jpg"><source src="https://cdn.ainative.foundation/video/20260915_1a82de6b6ee541bbb18eeb9c3a9c8ab8.mp4" type="video/mp4"></video><br />
Video Credit: @OpenAIDevs on X</p>
<div style="width:100%;height:2px;background:#808080;margin:10px 0"></div>
<p>That&#8217;s all for today&#8217;s Global AI Native Industry Insights. Join us at <a href="https://member.ainativefoundation.org/">AI Native Foundation Membership Dashboard</a> for the latest insights on AI Native, or follow our linkedin account at <a href="https://www.linkedin.com/company/ainativefoundation/">AI Native Foundation</a> and our twitter account at <a href="https://x.com/AINativeF">AINativeF</a>.</p>
<p>The post <a href="https://ainativefoundation.org/global-ai-native-industry-insights-20260915-elevenlabs-perplexity-ai-cognition-more/">Global AI Native Industry Insights &#8211; 20260915 &#8211;  ElevenLabs | Perplexity AI | Cognition | more</a> appeared first on <a href="https://ainativefoundation.org">AI Native Foundation</a>.</p>
]]></content:encoded>
					
		
		<enclosure url="https://cdn.ainative.foundation/video/20260915_ad60cec8e60d45d8b73aacce38f4008c.mp4" length="17081978" type="video/mp4" />
<enclosure url="https://cdn.ainative.foundation/video/20260915_7d18bfd304ad440aa0e3bbfd53195a53.mp4" length="4151936" type="video/mp4" />
<enclosure url="https://cdn.ainative.foundation/video/20260915_gj_video_aicoding.mp4" length="6106418" type="video/mp4" />
<enclosure url="https://cdn.ainative.foundation/video/20260915_1a82de6b6ee541bbb18eeb9c3a9c8ab8.mp4" length="6937674" type="video/mp4" />

			</item>
	</channel>
</rss>
