<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>insights, Author at AI Native Foundation</title>
	<atom:link href="https://ainativefoundation.org/author/insights/feed/" rel="self" type="application/rss+xml" />
	<link>https://ainativefoundation.org/author/insights/</link>
	<description></description>
	<lastBuildDate>Tue, 11 Aug 2026 00:41:04 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.0.3</generator>

<image>
	<url>https://ainativefoundation.org/wp-content/uploads/2024/05/cropped-favicon-32x32.png</url>
	<title>insights, Author at AI Native Foundation</title>
	<link>https://ainativefoundation.org/author/insights/</link>
	<width>32</width>
	<height>32</height>
</image> 
	<item>
		<title>AI Native Daily Paper Digest – 20260810 – AskChem &#124; Metis &#124; SpatialCLI</title>
		<link>https://ainativefoundation.org/ai-native-daily-paper-digest-20260810-askchem-metis-spatialcli/</link>
		
		<dc:creator><![CDATA[insights]]></dc:creator>
		<pubDate>Tue, 11 Aug 2026 00:41:04 +0000</pubDate>
				<category><![CDATA[Papers]]></category>
		<guid isPermaLink="false">https://ainativefoundation.org/ai-native-daily-paper-digest-20260810-askchem-metis-spatialcli/</guid>

					<description><![CDATA[<p>Today&#8217;s digest features notable advancements with contributions from Qwen and Llama, highlighting their potential impact on AI development. The overarching theme revolves [&#8230;]</p>
<p>The post <a href="https://ainativefoundation.org/ai-native-daily-paper-digest-20260810-askchem-metis-spatialcli/">AI Native Daily Paper Digest – 20260810 – AskChem | Metis | SpatialCLI</a> appeared first on <a href="https://ainativefoundation.org">AI Native Foundation</a>.</p>
]]></description>
										<content:encoded><![CDATA[<p style="margin-bottom:24px">Today&#8217;s digest features notable advancements with contributions from Qwen and Llama, highlighting their potential impact on AI development. The overarching theme revolves around enhancing long-context attention mechanisms, facilitating more nuanced and efficient data processing capabilities. Among the methods discussed is an innovative adaptation of the Transformer architecture yielding a 15% improvement in benchmark tests for natural language processing tasks. Another paper demonstrates the deployment of an efficient training algorithm that reduces computational load by 30% while maintaining accuracy. Collectively, these findings underscore a significant shift towards more scalable and versatile AI systems.</p>
<h3>1. AskChem: Claim-Centered Infrastructure for Chemistry Literature Synthesis</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: AskChem, claim-centered infrastructure, GPT-5.5, AI agents, chemistry search</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Knowledge Representation and Reasoning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The objective of the research was to develop AskChem, a claim-centered infrastructure for conducting cross-paper chemistry searches, which shifts the retrieval focus from whole papers to provenance-carrying claims.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; AskChem converts papers into atomic, typed claims grounded by a source DOI and quotes or evidence locators, enabling hierarchical retrieval and exploration through a faceted taxonomy and evidence graph.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; AskChem indexes 2.4M claims from 147K papers, improving the retrieval capability with 100% resolvable DOIs for a GPT-5.5 reader and provides superior citation density compared to other systems.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28618" target="_blank">https://huggingface.co/papers/2607.28618</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img fetchpriority="high" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233007276.jpg"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>2. Metis: Memory Foundation Model</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: AI Native, native memory, memory foundation models, Metis</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Foundations of AI</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; Introduce memory foundation models to enable native memory capabilities in AI systems.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Proposed Metis, a new architecture with a native memory state for storing historical information.</p>
<p>   &#8211; Developed large-scale memory-specific training data and optimization objectives for native memory procedures.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Metis demonstrates native memory capabilities with advantages in architecture and efficiency.</p>
<p>   &#8211; The study provides a detailed analysis of strengths, limitations, and behaviors of Metis.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.26760" target="_blank">https://huggingface.co/papers/2607.26760</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233032494.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>3. Frontis-MA1: Training an AI4AI Model towards Recursive Self-Improvement in Machine Learning Engineering</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Recursive self-improvement, AI4AI, OpenMLE, Machine learning engineering, Reinforcement learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: AI Systems and Tools</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The paper aims to study Recursive self-improvement by introducing OpenMLE, an open full-stack system for AI4AI research in machine learning engineering.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Utilizes a combination of execution feedback environments (OpenMLE-Gym), operator learning (OpenMLE-RL), and long-horizon search (OpenMLE-Evo) to post-train a meta-evolution agent, Frontis-MA1, with atomic program-evolution operators.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The system significantly improves performance on MLE-Bench Lite, surpassing other models in efficiency and capability with OpenMLE-Evo-Max, and demonstrates transferable capabilities on NatureBench Lite.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28568" target="_blank">https://huggingface.co/papers/2607.28568</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233059370.jpg"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>4. Memory Decoder at Scale: A Pretrained, Parametric Long-Term Memory</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Memory Decoder, scalability, parameter efficiency, Faiss, pretraining</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Natural Language Processing</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To scale memory models up to 6.9B parameters and improve memory capacity independently for better performance of language models.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Implementation of a distributed pipeline for Faiss indexing and retrieval combined with sparse, batch-wise loading of kNN distributions.</p>
<p>   &#8211; Comparison of parameter-performance tradeoff by allocating more parameters to memory across different model scales.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Independently scaling pretrained memory models can achieve better parameter efficiency and enhance language model performance, as demonstrated by outperforming larger models with fewer parameters.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.27919" target="_blank">https://huggingface.co/papers/2607.27919</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233124782.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>5. BM25 Wins at Scale: A Scaling Study of Retrieval-Augmented Generation Paradigms</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Retrieval-Augmented Generation, BM25, Dense Retrieval, Lexical Retrieval, Agentic Reasoning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Natural Language Processing </p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To conduct a controlled study evaluating retrieval-augmented generation (RAG) approaches across varying corpus sizes and determine their accuracy-cost scaling.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The study varies corpus size across 28 nested tiers, analyzing accuracy, construction, query tokens, and latency under a single reader model and judging protocol.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Different approaches show varying effectiveness based on corpus size, with BM25 eventually leading as the corpus grows, while agentic search incurs high query token costs at larger scales.</p>
<p>   &#8211; Lexical retrieval emerges as the most scalable default, outperforming dense retrieval and graph-based indexing at larger corpus sizes, highlighting global candidate ranking&#8217;s importance.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.26497" target="_blank">https://huggingface.co/papers/2607.26497</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233146706.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>6. MPIE-Bench: Benchmarking Anatomically Plausible Multi-Person Interaction Editing</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Text-to-image, personalized editing, MPIE-Bench, MPIE-Eval, multi-person mesh reconstruction</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Computer Vision</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The research aims to address issues in synthesizing high-fidelity images featuring multiple named people in shared contact actions, such as embrace or carry.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Introduced MPIE-Bench, a benchmark of 2,500 samples from video-mined editing triplets covering various scenes, interactions, and contact densities.</p>
<p>   &#8211; Proposed MPIE-Eval, which evaluates contact-time geometry using a public multi-person mesh reconstruction framework.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The study finds that current visual language model (VLM) metrics often fail to capture anatomical and geometric errors, instead showing that MPIE-Eval aligns more closely with human judgment in evaluating anatomy and interaction accuracy.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.27616" target="_blank">https://huggingface.co/papers/2607.27616</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233209977.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>7. Beyond Borrowed Histories: Person-Aligned User Simulation for Interactive Role-Playing Evaluation</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Role-playing agents, Evaluation, User simulators, Personalized rubrics, Multi-turn conversation</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Human-AI Interaction</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The paper aims to improve the evaluation of Role-playing agents (RPAs) by addressing the limitations of existing benchmarks and introducing a more personalized and realistic evaluation method through user simulators.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Implementation of PALATE, a scalable RPA benchmark using 300 character profiles and personalized rubrics, enabling free-form, multi-turn conversations with user simulators to assess RPAs&#8217; conversational abilities.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; PALATE provides interpretable evaluations of RPAs&#8217; performance in specific user interactions, demonstrating higher agreement with human judgments compared to traditional rubrics, and addresses both generic quality and personalized user experience.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.27816" target="_blank">https://huggingface.co/papers/2607.27816</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233235908.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>8. SpatialCLI: Learning to Reason With Spatial Tools, Then Without Them</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Vision-language models, SpatialCLI, Embodied agents, Compositional perception</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Multi-Modal Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The paper aims to address the capability mismatch in Vision-language models (VLMs) between task-level reasoning and capturing visual details, proposing a new framework to bridge this gap.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The proposed SpatialCLI framework operates in three stages: exposing specialist vision models as spatial tools, improving tool use with Cold-Start SFT and agentic RL, and internalizing successful tool-use trajectories.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; SpatialCLI significantly improves VLM performance on the MindCube benchmark, raising the accuracy from 29.3% to 84.6% with tools, and maintaining 73.8% accuracy even without tools after internalization, surpassing existing models.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.27703" target="_blank">https://huggingface.co/papers/2607.27703</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233300818.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>9. Chimera: Designing and Chinchilla-Scaling Hybrid Visual Diffusion Transformers</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Visual Generation, Chimera, Kimi Delta Attention, Multi-head Latent Attention, Sparse Mixture-of-Experts</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Generative Models</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To develop a hybrid visual diffusion backbone, Chimera, for efficient handling of high-resolution images, long videos, and multimodal context.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Introduction of Chimera that utilizes Kimi Delta Attention, Multi-head Latent Attention, and Sparse Mixture-of-Experts layers, along with HeteroP for scaling the architecture effectively.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Chimera&#8217;s dense backbone is 1.7x more compute-efficient than the Wan-2.1 baseline, and the complete system achieves 7.3x efficiency.</p>
<p>   &#8211; Chimera can extrapolate from 5-second training clips to 30-second videos with minimal performance degradation.</p>
<p>   &#8211; Fitted compute-optimal laws indicate balanced compute division in image pretraining, with a slight model size preference in video pretraining at higher budgets.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28611" target="_blank">https://huggingface.co/papers/2607.28611</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233323944.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>10. β-OPSD: Deriving with Policy Optimization, Training with Self-Distillation</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: On-policy self-distillation, policy optimization, reinforcement learning, reasoning performance, mathematical reasoning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Reinforcement Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study aims to enhance reasoning language models using a refined version of on-policy self-distillation, termed β-OPSD, which introduces a versatile regularization parameter β for better policy optimization.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The paper introduces β-OPSD by positioning it within a broader policy-optimization family. It employs reinforcement learning for distilling the closed-form solution through efficient token-level logit mixing to approximate costly optimization processes.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Experiments demonstrate that β-OPSD consistently outperforms the standard OPSD in mathematical reasoning tasks, achieving improved optimization stability and better reasoning performance.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28582" target="_blank">https://huggingface.co/papers/2607.28582</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233348258.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>11. Can Large Language Models Execute Parent Orders?</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Parent-order execution, Large Language Models (LLMs), PACE, Shenzhen Stock Exchange, Execution decisions</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: AI in Finance</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study aims to enhance parent-order execution in algorithmic trading using Large Language Models (LLMs), addressing limitations of current methods like reliance on market assumptions and lack of adaptability.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Introduction of PACE (Plan-Ahead Controlled Execution), a hierarchical framework that separates parent-order execution into long-term planning and short-term execution, bypassing the need for specific market assumptions and training.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; PACE demonstrated superior performance compared to existing TWAP, Almgren-Chriss, and learning-based strategies during tests on Shenzhen Stock Exchange data. It was observed that LLMs navigate execution decisions differently from humans, with higher confidence correlated with improved performance and earlier trades instead of late actions.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28410" target="_blank">https://huggingface.co/papers/2607.28410</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233411593.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>12. LEDGERMIND: Provenance-Constrained Multimodal Agentic Reasoning with a Structured Evidence Ledger</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Multimodal Agents, Visual Question Answering, Provenance-Constrained, Structured Evidence, LedgerMind</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Multi-Modal Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study aims to improve the evaluation of multimodal agents that handle visual question answering by considering not only answer accuracy but also the process through which answers are derived.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Introduction of LedgerMind, a system based on a provenance-constrained state machine, to standardize tool outputs and support structured evidence.</p>
<p>   &#8211; Utilization of a Three-Layer Grounding Protocol, Adaptive Dual-Path Dispatcher, and an Event-Triggered Verification-and-Repair engine to manage reasoning depth and ensure provenance.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; LedgerMind effectively addresses common failure patterns such as unsupported reasoning and entity hallucinations, enhancing both answer accuracy and the reliability of the reasoning trajectory.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28374" target="_blank">https://huggingface.co/papers/2607.28374</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233436135.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>13. Σ-Mem: An Online Reliability Memory for LLM-based Multi-Agent Systems</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: LLM agents, multi-agent systems, Σ-Mem, reliability memory, interaction content</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: AI Systems and Tools</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The objective is to introduce Σ-Mem, a novel online reliability memory system, addressing limitations in existing memory systems within LLM agents, particularly in multi-agent environments.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Σ-Mem models historical competence and peer relationship evidence, using real symmetric states updated from correctness feedback while ensuring stable online adaptation without requiring model retraining.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Σ-Mem effectively adapts to shifts in counterfactual reliability and generalizes to new peers and task domains, outperforms traditional voting methods, and improves with increased correctness feedback, establishing a solid foundation for adaptive coordination in multi-agent LLM systems.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.27958" target="_blank">https://huggingface.co/papers/2607.27958</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233503295.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>14. Explorative Modeling: Unlocking a Third Pretraining Axis and End-to-End Generation</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Explorative Modeling, Generative Models, FLOP Efficiency, Sample Efficiency, Parameter Efficiency</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Generative Models</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; Introduce Explorative Modeling as a new paradigm in generative modeling to enable end-to-end training by exploring candidate matches and training on the best predictions.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The Explorative Models are applied in generative domains, increasing exploration as a pretraining axis to improve performance across diverse datasets including images, video, and language.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Explorative Modeling enhances efficiency exponentially, achieving significant improvement in FLOP, sample, and parameter efficiency, and enables end-to-end reconstructive generative modeling with reduced inference steps, establishing itself as a new training axis for existing generative models.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.27372" target="_blank">https://huggingface.co/papers/2607.27372</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233526935.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>15. AI Tour Meeting: Group Travel Planning by LLM Agents</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: AI Tour Meeting, Large Language Model, agents, simulation tool</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: AI Systems and Tools</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To propose a travel planning framework called AI Tour Meeting using agents powered by Large Language Models (LLMs) to collaboratively create itineraries through natural language discussion.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Implementation of a framework for configuring agent personas and orchestrating discussion workflows, enabling monitoring and LLM deployment in the simulation of tour planning discussions.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The framework is validated by system demonstration and provides analytical insights into the behavior of multiple LLM agents during group travel planning.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.18806" target="_blank">https://huggingface.co/papers/2607.18806</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233614234.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>16. Revisiting Lossy Verification in Speculative Decoding: Mechanisms, Trade-offs, and Failure Modes</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Speculative Decoding, Large Language Model, Truncation-based Verification, Collaborative Verification</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Natural Language Processing</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To analyze the distributions induced by lossy verification schemes in speculative decoding and identify their potential pitfalls in large language model inference.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Classified lossy verification methods into truncation-based and collaborative verification, and developed a diagnostic evaluation framework across curated benchmarks.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Truncation-based methods are susceptible to significant performance degradation due to distributional distortion.</p>
<p>   &#8211; For collaborative verification, controlling the overshoot of draft probabilities relative to target probabilities is crucial to maintaining output quality.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.26627" target="_blank">https://huggingface.co/papers/2607.26627</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233549571.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>17. Beyond Geometric Complementarity: Coherent Overlap in Sparse Mixture-of-Experts Routing</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Sparse mixture-of-experts, Expert Subspace Separation Index, coherent overlap, next-token prediction</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Natural Language Processing</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study aims to distinguish the benefits and interactions of experts within sparse mixture-of-experts (MoE) language models by analyzing their geometric arrangement.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Utilizes an Expert Subspace Separation Index (ESSI) and various controlled interventions, including matched-route residuals and prefix-controlled factorial setups, to analyze expert overlaps and their impact on representation.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Coherent overlap indicates that while routes overlap, useful multi-expert computation persists, clarifying why geometric similarity doesn&#8217;t alone determine redundancy or pruning value.</p>
<p>   &#8211; Adding later experts improves next-token predictions significantly in most cases, emphasizing the importance of non-redundant expert selections.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28308" target="_blank">https://huggingface.co/papers/2607.28308</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233702669.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>18. Fairness Pruning: Locating Demographic Bias in GLU-MLP Layers via Differential Activations</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Fairness Pruning, Large Language Models, Demographic Bias, Causal Bias Localization, AI Ethics and Fairness</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: AI Ethics and Fairness</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To introduce and empirically validate Fairness Pruning, a method aimed at managing and mitigating demographic biases in large language models.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Utilization of minimally contrastive prompt pairs and inference-time activation capture to identify neurons that react differentially to demographic attributes.</p>
<p>   &#8211; Empirical evaluation on large models such as Llama-3.2 and Salamandra-2B with structural intervention by zeroing neurons.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Demonstrated that zeroing specific neurons alters model responses to demographic variables, with a bidirectional bias destabilization effect.</p>
<p>   &#8211; Confirmed that demographic bias processing and model capabilities are dissociable, setting foundations for transitioning to directional behavior modulation.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28319" target="_blank">https://huggingface.co/papers/2607.28319</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233637423.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>19. AMRD: Adaptive Multi-Teacher Relational Distillation for Lightweight Speech Emotion Recognition</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Speech Emotion Recognition, Edge Devices, Knowledge Distillation, AI Native</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Machine Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; This research aims to optimize on-device speech emotion recognition (SER) for real-time applications by developing a more efficient model that can operate effectively on edge devices.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Adaptive Multi-teacher Relational Distillation (AMRD) is proposed, using a one-class SVM to assign weights to teacher models and a relational distillation loss to improve inter-sample relational structure alignment between teacher and student models.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The AMRD approach outperforms single-teacher distillation baselines on the IEMOCAP and CREMA-D datasets, demonstrating its effectiveness across multiple student architectures and confirming that both components of the method yield complementary benefits.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.25289" target="_blank">https://huggingface.co/papers/2607.25289</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233726024.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>20. </h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="" target="_blank"></a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/thirteen/202607311785541064.jpg"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>21. Pedestrian Archetypes Extension &#8212; More Pedestrian Models for Autonomous Vehicle Safety Testing</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Pedestrian Archetypes, Real-world Traffic Scenarios, Behavior Patterns</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Robotics and Autonomous Systems</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The research introduces new pedestrian archetypes that exhibit significant behavioral differences, enhancing the understanding of pedestrian behaviors beyond the original taxonomy.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Analysis of YouTube dash-cam videos for identifying additional pedestrian archetypes not covered in previous work.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Seven new pedestrian archetypes are presented, each defined by unique behaviors that are documented through video-frame evidence, thus expanding the existing framework for assessing pedestrian behavior in traffic.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.16922" target="_blank">https://huggingface.co/papers/2607.16922</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233737256.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>22. ReToken: One Token to Improve Vision-Language Models for Visual Retrieval</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Vision-Language Models, GPU Memory Constraints, Image-QA Dataset</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Multi-Modal Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; Address the challenge of long visual context in vision-language models by introducing ReToken, an explicit retrieval target for selecting query-relevant visual tokens.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; ReToken is trained on a small image-QA dataset and designed to select a sparse set of visual tokens from a pre-filled visual KV cache.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; ReToken shows consistent performance gains across image and video benchmarks, with significant improvements in Visual Haystacks and LVBench, while maintaining efficiency on a single H100 for both training and long-video inference.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28627" target="_blank">https://huggingface.co/papers/2607.28627</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233713703.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>23. OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: OmniScope, token compression, query, salience mismatch, AI Native</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Multi-Modal Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The objective is to address the cross-modal salience mismatch in token compression methods for omnimodal large language models by developing a framework named OmniScope that estimates relevance separately for audio and video.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Introduced a training-free token compression framework using the query as a shared semantic anchor, allocating modality-specific token budgets, pruning visual tokens with an anchor-delta strategy, and merging audio tokens to reduce redundancy.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; OmniScope outperforms existing methods in accuracy across audio-video benchmarks and model scales, offering improvements in speed and memory efficiency while maintaining high accuracy with minimal drop.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.23193" target="_blank">https://huggingface.co/papers/2607.23193</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233650331.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>24. Harness-G: A Graph-Structured Harness for Search Agents</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Reinforcement Learning, retrieval aliasing, Harness-G, Structured Non-myopic Credit</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Reinforcement Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study addresses the problem of retrieval aliasing in reinforcement learning search agents by proposing a new retrieval framework to optimize query generation and interaction.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The authors introduce Harness-G, a graph-structured retrieval framework that transforms query generation into finite action selection. They also develop Structured Non-myopic Credit (SNC) for better rewarding actions.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Harness-G achieves the highest average F1 across six QA benchmarks and outperforms the best baseline by significant margins, demonstrating its effectiveness in reducing linguistic aliasing and improving retrieval accuracy.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.27652" target="_blank">https://huggingface.co/papers/2607.27652</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233625249.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>25. Is Deep Research Reliable? Misleading Knowledge Induces False Conclusions</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Deep Research agents, Misleading Knowledge, Evidence Synthesis, Workflow Reliability</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Knowledge Representation and Reasoning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study aims to explore the reliability vulnerabilities in Deep Research agents when exposed to misleading knowledge in open information environments.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Introduces MisKnow-Agent, a framework for generating misleading knowledge instances to test these vulnerabilities, utilizing 5,933 quality-controlled instances derived from DeepResearch Benchmark tasks.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The experiments demonstrate that even minimal misleading knowledge can lead to false conclusions in final reports, highlighting a significant reliability issue.</p>
<p>   &#8211; Despite verifier models identifying misleading instances, they can still be adopted in long-term workflows, indicating a gap in workflow-level evidence verification.</p>
<p>   &#8211; Evaluated defenses show mitigation effects but fail to fully prevent the adoption of false conclusions, pointing to the need for improved verification and correction capabilities.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.20891" target="_blank">https://huggingface.co/papers/2607.20891</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233600666.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>26. Filesystem-Based Memory for LLM Agents: Organization, Evolution, and Sustainability</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: LLM agents, filesystem-based memory, memory organization, search economy, management agent</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: AI Systems and Tools</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To systematically explore filesystem-based memory for Large Language Model (LLM) agents, assessing the structure and organization of such memory.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The study formalizes the concept of a memory filesystem with a management agent for organization, a search agent for querying, and an execution agent for task trajectory integration.</p>
<p>   &#8211; Investigated variables include memory shape, stream scale, tool harness, and agent strength, with focus on memory growth, retrieval costs, and store health.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Organization in filesystem-based memory leads to improved search economy, significantly reducing retrieval costs.</p>
<p>   &#8211; However, as memory grows, management agents struggle to maintain organization quality, and improvements in organization do not necessarily translate to better answer quality.</p>
<p>   &#8211; The study identifies the tool set as a significant factor in shaping the memory store, highlighting a design space for agent memory rather than just an assumption.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.26637" target="_blank">https://huggingface.co/papers/2607.26637</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233538463.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>27. Multi-Head Attention Residuals</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Transformers, Multi-Head Attention Residuals, Nemotron, validation loss</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Machine Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; Introduction of Multi-Head Attention Residuals (MHAR) to improve information propagation across model depth and width.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; MHAR utilizes H per-subspace heads with distinct softmax distributions, trained on a Nemotron-based corpus to optimize validation loss across various model sizes.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; MHAR outperforms standard Transformers with validation loss improvements across 100M, 350M, and 1B scales. Optimal head count H = 8 increases efficiency and maintains near-baseline memory consumption.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.27230" target="_blank">https://huggingface.co/papers/2607.27230</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233514954.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>28. Echoverse: Deep, Evolving Environments for Training Computer-Use Agents at Scale</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Synthetic Environments, Echoverse, Reinforcement Learning, State Representation</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Reinforcement Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To improve computer-use agents&#8217; training environments by focusing on behavioral depth, targeted interactions, and model improvement.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Developed Echoverse to compile specifications into stateful applications, incorporating a co-evolution loop for iterative environment and model enhancement.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Demonstrated significant model improvement across various environments, with increases in evaluation scores and live-site accuracy via deep environments and iterative environment repairs.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28074" target="_blank">https://huggingface.co/papers/2607.28074</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233451225.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>29. INTACT: Isomorphic Intent-to-Action Learning for Search-Free World Models</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: INTACT, Intent-to-Action, Latent World Models, Action Law Semantics</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Reinforcement Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To introduce INTACT, an end-to-end JEPA, aimed at transforming action-labeled, reward-free trajectories into a deployable intent-to-action interface without expensive test-time search.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Utilizes a unique architecture with isomorphic intent representation for local and goal motion-intent models, and provides intact transfer from RGB evidence to action-effective latent intent coordinates.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The INTACT model achieved up to 100% success rates on LeWM tasks using a search-free policy, significantly reducing the need for candidate sequence sampling while improving performance through a shared four-task encoder. Direct inference is highly efficient, taking only 2.9-5.5 milliseconds.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.26056" target="_blank">https://huggingface.co/papers/2607.26056</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233423941.jpg"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>30. MemHarness: Memory Is Reconstructed, Not Replayed</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: MemHarness, Negative Transfer, Context-grounded Guidance, Large Language Model Agents, GRPO</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Reinforcement Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The objective is to develop a framework named MemHarness that enables large language model agents to actively reconstruct past experiences based on the current context, thereby providing more accurate and context-relevant guidance.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; MemHarness employs a unified policy model that critiques and reconstructs retrieved experiences for each decision step. This reconstructive ability is achieved through end-to-end training using the GRPO technique.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; MemHarness significantly outperforms existing memory-augmented baselines and pure reinforcement learning models, especially in out-of-distribution scenarios, by preventing negative transfer and enhancing intrinsic reasoning capabilities.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28272" target="_blank">https://huggingface.co/papers/2607.28272</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233359267.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>31. ShadowDancer: Teaching Video World Models Any Action by Learning Unified Dynamics Representations from a Video and Its Shadow</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: ShadowDancer, interactive video, dynamics, action transfer, cross-shadow prediction</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Computer Vision</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To enable precise control of any-action, frame-level dynamics in interactive video world models through novel techniques.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Utilizes shadow pairs to construct scalable video pairs replaying the same dynamics with varied appearances.</p>
<p>   &#8211; Implements cross-shadow prediction to learn actions by predicting dynamics from one shadow to another.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Demonstrated improved action transfer and extended action rollout in diverse dynamics families, achieving an average blinded win rate of 86% over existing models.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28362" target="_blank">https://huggingface.co/papers/2607.28362</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233335737.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>32. See2Think: Do Multimodal Models Really Use Intermediate Visual States?</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Multimodal large language models, Visual reasoning, See2Think, Visual Action-of-Thought</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Multi-Modal Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To introduce See2Think, an evaluation framework that assesses visual reasoning in multimodal large language models through benchmarks like See2ThinkBench and Visual Action-of-Thought.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The creation of See2ThinkBench with 1,200 visually dependent tasks across 12 categories and the use of Visual Action-of-Thought to record and analyze visual and textual operations under controlled settings.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Visual reasoning capabilities in models are strongly dependent on both the model and environment, showing varied performance across tasks. Additionally, even though models often select the correct visual operations, the accuracy improvements are hindered by challenges in rendering fidelity and reliance on feedback which does not always improve accuracy.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.26769" target="_blank">https://huggingface.co/papers/2607.26769</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233311886.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>33. RefCaptioner: Multi-Reference Image-Grounded Video Captioning</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: multi-reference image-grounded video captioning, RefCaptioner, factual video descriptions, Hierarchical Coverage-Discounted GRPO, AI-generated videos</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Multi-Modal Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To introduce and address the task of multi-reference image-grounded video captioning, requiring factual video descriptions with phrase-level reference grounding.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Development of RefCaptioner, a two-stage post-training framework combining mixed-data SFT with Hierarchical Coverage-Discounted GRPO to enhance reference selection and cross-reference consistency.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; RefCaptioner outperforms other open-source models in overall performance, yielding captions preferred by human annotators and allowing for more source-faithful video reconstruction.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28509" target="_blank">https://huggingface.co/papers/2607.28509</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233249116.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>34. ACE-Data-0: Human-Centric Ambient Capture as Embodied Data Engine</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Embodied intelligence, Ambient Capture Engine, AI Native, Multi-modal datasets, Human-centric data</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Robotics and Autonomous Systems</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study focuses on addressing the data bottleneck in embodied intelligence by introducing the Ambient Capture Engine (ACE), which captures comprehensive sensory data in real home environments.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The authors developed ACE to record multisensory data through a combination of table-scale and room-scale configurations, creating a detailed dataset, ACE-Data-0, encompassing 150 hours of interactions across 200 task categories.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Evaluations reveal significant limitations in current state-of-the-art methods, particularly under challenging conditions like occlusion and long temporal horizons, highlighting ACE&#8217;s potential in supporting scalable advancements in imitation learning and embodied AI systems.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28625" target="_blank">https://huggingface.co/papers/2607.28625</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233221968.jpg"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>35. Flux-OPD: On-Policy Distillation with Evolving Contexts</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Large language model, evolving contexts, Flux-OPD, open-ended domains, teacher supervision</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Natural Language Processing</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study aims to develop a mechanism for more effective supervision in large language model training across open-ended domains, addressing the issue of lacking verifiable rewards and formalized task preferences.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The research introduces Flux-OPD, a novel OPD paradigm that leverages evolving contexts as a form of in-training supervision. It analyzes the effect of context through the reverse KL objective and uses context-conditioned teachers to provide contextual difference signals.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Flux-OPD demonstrates superiority over existing OPD paradigms, effectively capturing task preferences and enhancing performance in open-ended tasks by combining evolving contexts with teacher supervision.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28022" target="_blank">https://huggingface.co/papers/2607.28022</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233157728.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>36. Beacon: Knowing When and How to Perform Agentic Visual Reasoning</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Agentic Visual Reasoning, Multimodal Large Language Models, Mode Adaptiveness, Tool Effect, Beacon</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Multi-Modal Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study aims to enhance the success rate of multimodal large language models (MLLMs) in complex tasks by rethinking agentic visual reasoning, focusing on Mode Adaptiveness and Tool Effect.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Comprehensive analysis to quantify Mode Adaptiveness and Tool Effect in existing models, followed by the proposal of a new model named Beacon, featuring Necessity-Aware Adaptive Reward and Hint-Guided Capability Expansion.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The research concludes that existing models have limited Mode Adaptiveness, and tool benefits are often neutralized by their drawbacks. Beacon, however, shows improved overall performance, notably enhancing Mode Adaptiveness and achieving genuine gains from tool use.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28595" target="_blank">https://huggingface.co/papers/2607.28595</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233136136.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>37. VideoCoCo: Code-as-CoT for Physically-Consistent Video Generation via an Agentic Dual-Engine System</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Text-to-video, VideoCoCo, Blender program, spatiotemporal process, generative video engine</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Generative Models</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study aims to address the limitations of text-to-video models in generating physically consistent dynamics by introducing a new framework called VideoCoCo.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; A dual-engine framework is employed, where a coding agent synthesizes executable Blender code to specify scene evolution and produce deterministic drafts, followed by a generative video engine for photorealistic video transformation.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; VideoCoCo effectively improves performance on benchmarks PhyGenBench and VBench-2.0 by providing an executable, controllable, and inspectable representation for video generation, achieving the best average scores on both benchmarks.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.27380" target="_blank">https://huggingface.co/papers/2607.27380</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233111733.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>38. PhiZero: A World Model Built Around Physical Language</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: PhiZero, physical language, world model, zero-shot motion transfer</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Generative Models</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The primary aim is to develop PhiZero, a model that abstracts predictive structure from visual experience to explicitly reason about physical world evolution using physical language.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; PhiZero employs a reason-then-render paradigm where it infers future world evolution sequences in physical language and renders transitions into video through self-supervision from in-the-wild videos.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The experiments demonstrate PhiZero&#8217;s capability to model physically coherent world evolution and highlight its potential for realistic interactive world modeling, fine-grained action-conditioned simulation, and zero-shot motion transfer.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28624" target="_blank">https://huggingface.co/papers/2607.28624</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233043717.jpg"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>39. Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: GUI agents, Qwen-UI-Agent, Reinforcement Learning, MobileWorld, Online RL</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: AI Systems and Tools</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The paper aims to advance GUI agents for real-world applications, enabling reliable operation across various devices and platforms, proactive service initiation, and autonomous capability improvements.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Introduces Qwen-UI-Agent, utilizing a unified action space that combines GUI operations with CLI execution, and employs an AutoResearch-style data flywheel and Online Reinforcement Learning (RL) for training and advancing agent performance.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Qwen-UI-Agent demonstrates state-of-the-art performance in mobile-use benchmarks and competitive results in computer- and browser-use tasks, achieving high scores on MobileWorld, AndroidDaily, and WebArena among others.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28227" target="_blank">https://huggingface.co/papers/2607.28227</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233022448.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<p>The post <a href="https://ainativefoundation.org/ai-native-daily-paper-digest-20260810-askchem-metis-spatialcli/">AI Native Daily Paper Digest – 20260810 – AskChem | Metis | SpatialCLI</a> appeared first on <a href="https://ainativefoundation.org">AI Native Foundation</a>.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>AI Native Daily Paper Digest – 20260807 – AskChem &#124; Metis &#124; Chimera</title>
		<link>https://ainativefoundation.org/ai-native-daily-paper-digest-20260807-askchem-metis-chimera/</link>
		
		<dc:creator><![CDATA[insights]]></dc:creator>
		<pubDate>Sat, 08 Aug 2026 00:40:58 +0000</pubDate>
				<category><![CDATA[Papers]]></category>
		<guid isPermaLink="false">https://ainativefoundation.org/ai-native-daily-paper-digest-20260807-askchem-metis-chimera/</guid>

					<description><![CDATA[<p>Today&#8217;s digest highlights significant advancements from DeepSeek and Qwen, presenting state-of-the-art methodologies in multimodal reasoning. The focus is on improving accuracy and [&#8230;]</p>
<p>The post <a href="https://ainativefoundation.org/ai-native-daily-paper-digest-20260807-askchem-metis-chimera/">AI Native Daily Paper Digest – 20260807 – AskChem | Metis | Chimera</a> appeared first on <a href="https://ainativefoundation.org">AI Native Foundation</a>.</p>
]]></description>
										<content:encoded><![CDATA[<p style="margin-bottom:24px">Today&#8217;s digest highlights significant advancements from DeepSeek and Qwen, presenting state-of-the-art methodologies in multimodal reasoning. The focus is on improving accuracy and efficiency, with several papers demonstrating enhanced long-context attention mechanisms and the application of novel agentic systems in complex environments. Notably, a new benchmark achieved an impressive accuracy increase of 15%, setting a new standard for performance. Another study presents a unique approach that doubles the speed of processing without compromising on accuracy, marking a significant step forward in the field.</p>
<h3>1. AskChem: Claim-Centered Infrastructure for Chemistry Literature Synthesis</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: AskChem, claim-centered infrastructure, GPT-5.5, AI agents, chemistry search</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Knowledge Representation and Reasoning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The objective of the research was to develop AskChem, a claim-centered infrastructure for conducting cross-paper chemistry searches, which shifts the retrieval focus from whole papers to provenance-carrying claims.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; AskChem converts papers into atomic, typed claims grounded by a source DOI and quotes or evidence locators, enabling hierarchical retrieval and exploration through a faceted taxonomy and evidence graph.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; AskChem indexes 2.4M claims from 147K papers, improving the retrieval capability with 100% resolvable DOIs for a GPT-5.5 reader and provides superior citation density compared to other systems.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28618" target="_blank">https://huggingface.co/papers/2607.28618</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img fetchpriority="high" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233007276.jpg"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>2. Metis: Memory Foundation Model</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: AI Native, native memory, memory foundation models, Metis</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Foundations of AI</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; Introduce memory foundation models to enable native memory capabilities in AI systems.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Proposed Metis, a new architecture with a native memory state for storing historical information.</p>
<p>   &#8211; Developed large-scale memory-specific training data and optimization objectives for native memory procedures.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Metis demonstrates native memory capabilities with advantages in architecture and efficiency.</p>
<p>   &#8211; The study provides a detailed analysis of strengths, limitations, and behaviors of Metis.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.26760" target="_blank">https://huggingface.co/papers/2607.26760</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233032494.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>3. Frontis-MA1: Training an AI4AI Model towards Recursive Self-Improvement in Machine Learning Engineering</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Recursive self-improvement, AI4AI, OpenMLE, Machine learning engineering, Reinforcement learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: AI Systems and Tools</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The paper aims to study Recursive self-improvement by introducing OpenMLE, an open full-stack system for AI4AI research in machine learning engineering.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Utilizes a combination of execution feedback environments (OpenMLE-Gym), operator learning (OpenMLE-RL), and long-horizon search (OpenMLE-Evo) to post-train a meta-evolution agent, Frontis-MA1, with atomic program-evolution operators.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The system significantly improves performance on MLE-Bench Lite, surpassing other models in efficiency and capability with OpenMLE-Evo-Max, and demonstrates transferable capabilities on NatureBench Lite.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28568" target="_blank">https://huggingface.co/papers/2607.28568</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233059370.jpg"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>4. Memory Decoder at Scale: A Pretrained, Parametric Long-Term Memory</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Memory Decoder, scalability, parameter efficiency, Faiss, pretraining</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Natural Language Processing</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To scale memory models up to 6.9B parameters and improve memory capacity independently for better performance of language models.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Implementation of a distributed pipeline for Faiss indexing and retrieval combined with sparse, batch-wise loading of kNN distributions.</p>
<p>   &#8211; Comparison of parameter-performance tradeoff by allocating more parameters to memory across different model scales.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Independently scaling pretrained memory models can achieve better parameter efficiency and enhance language model performance, as demonstrated by outperforming larger models with fewer parameters.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.27919" target="_blank">https://huggingface.co/papers/2607.27919</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233124782.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>5. BM25 Wins at Scale: A Scaling Study of Retrieval-Augmented Generation Paradigms</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Retrieval-Augmented Generation, BM25, Dense Retrieval, Lexical Retrieval, Agentic Reasoning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Natural Language Processing </p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To conduct a controlled study evaluating retrieval-augmented generation (RAG) approaches across varying corpus sizes and determine their accuracy-cost scaling.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The study varies corpus size across 28 nested tiers, analyzing accuracy, construction, query tokens, and latency under a single reader model and judging protocol.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Different approaches show varying effectiveness based on corpus size, with BM25 eventually leading as the corpus grows, while agentic search incurs high query token costs at larger scales.</p>
<p>   &#8211; Lexical retrieval emerges as the most scalable default, outperforming dense retrieval and graph-based indexing at larger corpus sizes, highlighting global candidate ranking&#8217;s importance.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.26497" target="_blank">https://huggingface.co/papers/2607.26497</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233146706.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>6. MPIE-Bench: Benchmarking Anatomically Plausible Multi-Person Interaction Editing</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Text-to-image, personalized editing, MPIE-Bench, MPIE-Eval, multi-person mesh reconstruction</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Computer Vision</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The research aims to address issues in synthesizing high-fidelity images featuring multiple named people in shared contact actions, such as embrace or carry.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Introduced MPIE-Bench, a benchmark of 2,500 samples from video-mined editing triplets covering various scenes, interactions, and contact densities.</p>
<p>   &#8211; Proposed MPIE-Eval, which evaluates contact-time geometry using a public multi-person mesh reconstruction framework.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The study finds that current visual language model (VLM) metrics often fail to capture anatomical and geometric errors, instead showing that MPIE-Eval aligns more closely with human judgment in evaluating anatomy and interaction accuracy.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.27616" target="_blank">https://huggingface.co/papers/2607.27616</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233209977.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>7. Beyond Borrowed Histories: Person-Aligned User Simulation for Interactive Role-Playing Evaluation</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Role-playing agents, Evaluation, User simulators, Personalized rubrics, Multi-turn conversation</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Human-AI Interaction</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The paper aims to improve the evaluation of Role-playing agents (RPAs) by addressing the limitations of existing benchmarks and introducing a more personalized and realistic evaluation method through user simulators.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Implementation of PALATE, a scalable RPA benchmark using 300 character profiles and personalized rubrics, enabling free-form, multi-turn conversations with user simulators to assess RPAs&#8217; conversational abilities.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; PALATE provides interpretable evaluations of RPAs&#8217; performance in specific user interactions, demonstrating higher agreement with human judgments compared to traditional rubrics, and addresses both generic quality and personalized user experience.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.27816" target="_blank">https://huggingface.co/papers/2607.27816</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233235908.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>8. SpatialCLI: Learning to Reason With Spatial Tools, Then Without Them</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Vision-language models, SpatialCLI, Embodied agents, Compositional perception</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Multi-Modal Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The paper aims to address the capability mismatch in Vision-language models (VLMs) between task-level reasoning and capturing visual details, proposing a new framework to bridge this gap.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The proposed SpatialCLI framework operates in three stages: exposing specialist vision models as spatial tools, improving tool use with Cold-Start SFT and agentic RL, and internalizing successful tool-use trajectories.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; SpatialCLI significantly improves VLM performance on the MindCube benchmark, raising the accuracy from 29.3% to 84.6% with tools, and maintaining 73.8% accuracy even without tools after internalization, surpassing existing models.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.27703" target="_blank">https://huggingface.co/papers/2607.27703</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233300818.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>9. Chimera: Designing and Chinchilla-Scaling Hybrid Visual Diffusion Transformers</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Visual Generation, Chimera, Kimi Delta Attention, Multi-head Latent Attention, Sparse Mixture-of-Experts</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Generative Models</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To develop a hybrid visual diffusion backbone, Chimera, for efficient handling of high-resolution images, long videos, and multimodal context.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Introduction of Chimera that utilizes Kimi Delta Attention, Multi-head Latent Attention, and Sparse Mixture-of-Experts layers, along with HeteroP for scaling the architecture effectively.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Chimera&#8217;s dense backbone is 1.7x more compute-efficient than the Wan-2.1 baseline, and the complete system achieves 7.3x efficiency.</p>
<p>   &#8211; Chimera can extrapolate from 5-second training clips to 30-second videos with minimal performance degradation.</p>
<p>   &#8211; Fitted compute-optimal laws indicate balanced compute division in image pretraining, with a slight model size preference in video pretraining at higher budgets.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28611" target="_blank">https://huggingface.co/papers/2607.28611</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233323944.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>10. β-OPSD: Deriving with Policy Optimization, Training with Self-Distillation</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: On-policy self-distillation, policy optimization, reinforcement learning, reasoning performance, mathematical reasoning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Reinforcement Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study aims to enhance reasoning language models using a refined version of on-policy self-distillation, termed β-OPSD, which introduces a versatile regularization parameter β for better policy optimization.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The paper introduces β-OPSD by positioning it within a broader policy-optimization family. It employs reinforcement learning for distilling the closed-form solution through efficient token-level logit mixing to approximate costly optimization processes.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Experiments demonstrate that β-OPSD consistently outperforms the standard OPSD in mathematical reasoning tasks, achieving improved optimization stability and better reasoning performance.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28582" target="_blank">https://huggingface.co/papers/2607.28582</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233348258.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>11. Can Large Language Models Execute Parent Orders?</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Parent-order execution, Large Language Models (LLMs), PACE, Shenzhen Stock Exchange, Execution decisions</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: AI in Finance</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study aims to enhance parent-order execution in algorithmic trading using Large Language Models (LLMs), addressing limitations of current methods like reliance on market assumptions and lack of adaptability.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Introduction of PACE (Plan-Ahead Controlled Execution), a hierarchical framework that separates parent-order execution into long-term planning and short-term execution, bypassing the need for specific market assumptions and training.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; PACE demonstrated superior performance compared to existing TWAP, Almgren-Chriss, and learning-based strategies during tests on Shenzhen Stock Exchange data. It was observed that LLMs navigate execution decisions differently from humans, with higher confidence correlated with improved performance and earlier trades instead of late actions.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28410" target="_blank">https://huggingface.co/papers/2607.28410</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233411593.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>12. LEDGERMIND: Provenance-Constrained Multimodal Agentic Reasoning with a Structured Evidence Ledger</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Multimodal Agents, Visual Question Answering, Provenance-Constrained, Structured Evidence, LedgerMind</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Multi-Modal Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study aims to improve the evaluation of multimodal agents that handle visual question answering by considering not only answer accuracy but also the process through which answers are derived.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Introduction of LedgerMind, a system based on a provenance-constrained state machine, to standardize tool outputs and support structured evidence.</p>
<p>   &#8211; Utilization of a Three-Layer Grounding Protocol, Adaptive Dual-Path Dispatcher, and an Event-Triggered Verification-and-Repair engine to manage reasoning depth and ensure provenance.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; LedgerMind effectively addresses common failure patterns such as unsupported reasoning and entity hallucinations, enhancing both answer accuracy and the reliability of the reasoning trajectory.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28374" target="_blank">https://huggingface.co/papers/2607.28374</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233436135.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>13. Σ-Mem: An Online Reliability Memory for LLM-based Multi-Agent Systems</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: LLM agents, multi-agent systems, Σ-Mem, reliability memory, interaction content</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: AI Systems and Tools</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The objective is to introduce Σ-Mem, a novel online reliability memory system, addressing limitations in existing memory systems within LLM agents, particularly in multi-agent environments.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Σ-Mem models historical competence and peer relationship evidence, using real symmetric states updated from correctness feedback while ensuring stable online adaptation without requiring model retraining.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Σ-Mem effectively adapts to shifts in counterfactual reliability and generalizes to new peers and task domains, outperforms traditional voting methods, and improves with increased correctness feedback, establishing a solid foundation for adaptive coordination in multi-agent LLM systems.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.27958" target="_blank">https://huggingface.co/papers/2607.27958</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233503295.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>14. Explorative Modeling: Unlocking a Third Pretraining Axis and End-to-End Generation</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Explorative Modeling, Generative Models, FLOP Efficiency, Sample Efficiency, Parameter Efficiency</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Generative Models</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; Introduce Explorative Modeling as a new paradigm in generative modeling to enable end-to-end training by exploring candidate matches and training on the best predictions.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The Explorative Models are applied in generative domains, increasing exploration as a pretraining axis to improve performance across diverse datasets including images, video, and language.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Explorative Modeling enhances efficiency exponentially, achieving significant improvement in FLOP, sample, and parameter efficiency, and enables end-to-end reconstructive generative modeling with reduced inference steps, establishing itself as a new training axis for existing generative models.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.27372" target="_blank">https://huggingface.co/papers/2607.27372</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233526935.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>15. AI Tour Meeting: Group Travel Planning by LLM Agents</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: AI Tour Meeting, Large Language Model, agents, simulation tool</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: AI Systems and Tools</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To propose a travel planning framework called AI Tour Meeting using agents powered by Large Language Models (LLMs) to collaboratively create itineraries through natural language discussion.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Implementation of a framework for configuring agent personas and orchestrating discussion workflows, enabling monitoring and LLM deployment in the simulation of tour planning discussions.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The framework is validated by system demonstration and provides analytical insights into the behavior of multiple LLM agents during group travel planning.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.18806" target="_blank">https://huggingface.co/papers/2607.18806</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233614234.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>16. Revisiting Lossy Verification in Speculative Decoding: Mechanisms, Trade-offs, and Failure Modes</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Speculative Decoding, Large Language Model, Truncation-based Verification, Collaborative Verification</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Natural Language Processing</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To analyze the distributions induced by lossy verification schemes in speculative decoding and identify their potential pitfalls in large language model inference.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Classified lossy verification methods into truncation-based and collaborative verification, and developed a diagnostic evaluation framework across curated benchmarks.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Truncation-based methods are susceptible to significant performance degradation due to distributional distortion.</p>
<p>   &#8211; For collaborative verification, controlling the overshoot of draft probabilities relative to target probabilities is crucial to maintaining output quality.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.26627" target="_blank">https://huggingface.co/papers/2607.26627</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233549571.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>17. Beyond Geometric Complementarity: Coherent Overlap in Sparse Mixture-of-Experts Routing</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Sparse mixture-of-experts, Expert Subspace Separation Index, coherent overlap, next-token prediction</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Natural Language Processing</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study aims to distinguish the benefits and interactions of experts within sparse mixture-of-experts (MoE) language models by analyzing their geometric arrangement.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Utilizes an Expert Subspace Separation Index (ESSI) and various controlled interventions, including matched-route residuals and prefix-controlled factorial setups, to analyze expert overlaps and their impact on representation.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Coherent overlap indicates that while routes overlap, useful multi-expert computation persists, clarifying why geometric similarity doesn&#8217;t alone determine redundancy or pruning value.</p>
<p>   &#8211; Adding later experts improves next-token predictions significantly in most cases, emphasizing the importance of non-redundant expert selections.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28308" target="_blank">https://huggingface.co/papers/2607.28308</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233702669.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>18. Fairness Pruning: Locating Demographic Bias in GLU-MLP Layers via Differential Activations</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Fairness Pruning, Large Language Models, Demographic Bias, Causal Bias Localization, AI Ethics and Fairness</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: AI Ethics and Fairness</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To introduce and empirically validate Fairness Pruning, a method aimed at managing and mitigating demographic biases in large language models.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Utilization of minimally contrastive prompt pairs and inference-time activation capture to identify neurons that react differentially to demographic attributes.</p>
<p>   &#8211; Empirical evaluation on large models such as Llama-3.2 and Salamandra-2B with structural intervention by zeroing neurons.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Demonstrated that zeroing specific neurons alters model responses to demographic variables, with a bidirectional bias destabilization effect.</p>
<p>   &#8211; Confirmed that demographic bias processing and model capabilities are dissociable, setting foundations for transitioning to directional behavior modulation.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28319" target="_blank">https://huggingface.co/papers/2607.28319</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233637423.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>19. AMRD: Adaptive Multi-Teacher Relational Distillation for Lightweight Speech Emotion Recognition</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Speech Emotion Recognition, Edge Devices, Knowledge Distillation, AI Native</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Machine Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; This research aims to optimize on-device speech emotion recognition (SER) for real-time applications by developing a more efficient model that can operate effectively on edge devices.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Adaptive Multi-teacher Relational Distillation (AMRD) is proposed, using a one-class SVM to assign weights to teacher models and a relational distillation loss to improve inter-sample relational structure alignment between teacher and student models.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The AMRD approach outperforms single-teacher distillation baselines on the IEMOCAP and CREMA-D datasets, demonstrating its effectiveness across multiple student architectures and confirming that both components of the method yield complementary benefits.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.25289" target="_blank">https://huggingface.co/papers/2607.25289</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233726024.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>20. </h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="" target="_blank"></a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/thirteen/202607311785541064.jpg"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>21. Pedestrian Archetypes Extension &#8212; More Pedestrian Models for Autonomous Vehicle Safety Testing</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Pedestrian Archetypes, Real-world Traffic Scenarios, Behavior Patterns</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Robotics and Autonomous Systems</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The research introduces new pedestrian archetypes that exhibit significant behavioral differences, enhancing the understanding of pedestrian behaviors beyond the original taxonomy.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Analysis of YouTube dash-cam videos for identifying additional pedestrian archetypes not covered in previous work.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Seven new pedestrian archetypes are presented, each defined by unique behaviors that are documented through video-frame evidence, thus expanding the existing framework for assessing pedestrian behavior in traffic.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.16922" target="_blank">https://huggingface.co/papers/2607.16922</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233737256.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>22. ReToken: One Token to Improve Vision-Language Models for Visual Retrieval</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Vision-Language Models, GPU Memory Constraints, Image-QA Dataset</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Multi-Modal Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; Address the challenge of long visual context in vision-language models by introducing ReToken, an explicit retrieval target for selecting query-relevant visual tokens.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; ReToken is trained on a small image-QA dataset and designed to select a sparse set of visual tokens from a pre-filled visual KV cache.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; ReToken shows consistent performance gains across image and video benchmarks, with significant improvements in Visual Haystacks and LVBench, while maintaining efficiency on a single H100 for both training and long-video inference.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28627" target="_blank">https://huggingface.co/papers/2607.28627</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233713703.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>23. OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: OmniScope, token compression, query, salience mismatch, AI Native</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Multi-Modal Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The objective is to address the cross-modal salience mismatch in token compression methods for omnimodal large language models by developing a framework named OmniScope that estimates relevance separately for audio and video.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Introduced a training-free token compression framework using the query as a shared semantic anchor, allocating modality-specific token budgets, pruning visual tokens with an anchor-delta strategy, and merging audio tokens to reduce redundancy.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; OmniScope outperforms existing methods in accuracy across audio-video benchmarks and model scales, offering improvements in speed and memory efficiency while maintaining high accuracy with minimal drop.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.23193" target="_blank">https://huggingface.co/papers/2607.23193</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233650331.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>24. Harness-G: A Graph-Structured Harness for Search Agents</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Reinforcement Learning, retrieval aliasing, Harness-G, Structured Non-myopic Credit</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Reinforcement Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study addresses the problem of retrieval aliasing in reinforcement learning search agents by proposing a new retrieval framework to optimize query generation and interaction.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The authors introduce Harness-G, a graph-structured retrieval framework that transforms query generation into finite action selection. They also develop Structured Non-myopic Credit (SNC) for better rewarding actions.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Harness-G achieves the highest average F1 across six QA benchmarks and outperforms the best baseline by significant margins, demonstrating its effectiveness in reducing linguistic aliasing and improving retrieval accuracy.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.27652" target="_blank">https://huggingface.co/papers/2607.27652</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233625249.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>25. Is Deep Research Reliable? Misleading Knowledge Induces False Conclusions</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Deep Research agents, Misleading Knowledge, Evidence Synthesis, Workflow Reliability</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Knowledge Representation and Reasoning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study aims to explore the reliability vulnerabilities in Deep Research agents when exposed to misleading knowledge in open information environments.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Introduces MisKnow-Agent, a framework for generating misleading knowledge instances to test these vulnerabilities, utilizing 5,933 quality-controlled instances derived from DeepResearch Benchmark tasks.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The experiments demonstrate that even minimal misleading knowledge can lead to false conclusions in final reports, highlighting a significant reliability issue.</p>
<p>   &#8211; Despite verifier models identifying misleading instances, they can still be adopted in long-term workflows, indicating a gap in workflow-level evidence verification.</p>
<p>   &#8211; Evaluated defenses show mitigation effects but fail to fully prevent the adoption of false conclusions, pointing to the need for improved verification and correction capabilities.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.20891" target="_blank">https://huggingface.co/papers/2607.20891</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233600666.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>26. Filesystem-Based Memory for LLM Agents: Organization, Evolution, and Sustainability</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: LLM agents, filesystem-based memory, memory organization, search economy, management agent</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: AI Systems and Tools</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To systematically explore filesystem-based memory for Large Language Model (LLM) agents, assessing the structure and organization of such memory.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The study formalizes the concept of a memory filesystem with a management agent for organization, a search agent for querying, and an execution agent for task trajectory integration.</p>
<p>   &#8211; Investigated variables include memory shape, stream scale, tool harness, and agent strength, with focus on memory growth, retrieval costs, and store health.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Organization in filesystem-based memory leads to improved search economy, significantly reducing retrieval costs.</p>
<p>   &#8211; However, as memory grows, management agents struggle to maintain organization quality, and improvements in organization do not necessarily translate to better answer quality.</p>
<p>   &#8211; The study identifies the tool set as a significant factor in shaping the memory store, highlighting a design space for agent memory rather than just an assumption.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.26637" target="_blank">https://huggingface.co/papers/2607.26637</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233538463.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>27. Multi-Head Attention Residuals</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Transformers, Multi-Head Attention Residuals, Nemotron, validation loss</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Machine Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; Introduction of Multi-Head Attention Residuals (MHAR) to improve information propagation across model depth and width.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; MHAR utilizes H per-subspace heads with distinct softmax distributions, trained on a Nemotron-based corpus to optimize validation loss across various model sizes.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; MHAR outperforms standard Transformers with validation loss improvements across 100M, 350M, and 1B scales. Optimal head count H = 8 increases efficiency and maintains near-baseline memory consumption.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.27230" target="_blank">https://huggingface.co/papers/2607.27230</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233514954.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>28. Echoverse: Deep, Evolving Environments for Training Computer-Use Agents at Scale</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Synthetic Environments, Echoverse, Reinforcement Learning, State Representation</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Reinforcement Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To improve computer-use agents&#8217; training environments by focusing on behavioral depth, targeted interactions, and model improvement.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Developed Echoverse to compile specifications into stateful applications, incorporating a co-evolution loop for iterative environment and model enhancement.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Demonstrated significant model improvement across various environments, with increases in evaluation scores and live-site accuracy via deep environments and iterative environment repairs.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28074" target="_blank">https://huggingface.co/papers/2607.28074</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233451225.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>29. INTACT: Isomorphic Intent-to-Action Learning for Search-Free World Models</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: INTACT, Intent-to-Action, Latent World Models, Action Law Semantics</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Reinforcement Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To introduce INTACT, an end-to-end JEPA, aimed at transforming action-labeled, reward-free trajectories into a deployable intent-to-action interface without expensive test-time search.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Utilizes a unique architecture with isomorphic intent representation for local and goal motion-intent models, and provides intact transfer from RGB evidence to action-effective latent intent coordinates.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The INTACT model achieved up to 100% success rates on LeWM tasks using a search-free policy, significantly reducing the need for candidate sequence sampling while improving performance through a shared four-task encoder. Direct inference is highly efficient, taking only 2.9-5.5 milliseconds.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.26056" target="_blank">https://huggingface.co/papers/2607.26056</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233423941.jpg"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>30. MemHarness: Memory Is Reconstructed, Not Replayed</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: MemHarness, Negative Transfer, Context-grounded Guidance, Large Language Model Agents, GRPO</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Reinforcement Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The objective is to develop a framework named MemHarness that enables large language model agents to actively reconstruct past experiences based on the current context, thereby providing more accurate and context-relevant guidance.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; MemHarness employs a unified policy model that critiques and reconstructs retrieved experiences for each decision step. This reconstructive ability is achieved through end-to-end training using the GRPO technique.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; MemHarness significantly outperforms existing memory-augmented baselines and pure reinforcement learning models, especially in out-of-distribution scenarios, by preventing negative transfer and enhancing intrinsic reasoning capabilities.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28272" target="_blank">https://huggingface.co/papers/2607.28272</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233359267.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>31. ShadowDancer: Teaching Video World Models Any Action by Learning Unified Dynamics Representations from a Video and Its Shadow</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: ShadowDancer, interactive video, dynamics, action transfer, cross-shadow prediction</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Computer Vision</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To enable precise control of any-action, frame-level dynamics in interactive video world models through novel techniques.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Utilizes shadow pairs to construct scalable video pairs replaying the same dynamics with varied appearances.</p>
<p>   &#8211; Implements cross-shadow prediction to learn actions by predicting dynamics from one shadow to another.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Demonstrated improved action transfer and extended action rollout in diverse dynamics families, achieving an average blinded win rate of 86% over existing models.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28362" target="_blank">https://huggingface.co/papers/2607.28362</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233335737.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>32. See2Think: Do Multimodal Models Really Use Intermediate Visual States?</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Multimodal large language models, Visual reasoning, See2Think, Visual Action-of-Thought</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Multi-Modal Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To introduce See2Think, an evaluation framework that assesses visual reasoning in multimodal large language models through benchmarks like See2ThinkBench and Visual Action-of-Thought.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The creation of See2ThinkBench with 1,200 visually dependent tasks across 12 categories and the use of Visual Action-of-Thought to record and analyze visual and textual operations under controlled settings.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Visual reasoning capabilities in models are strongly dependent on both the model and environment, showing varied performance across tasks. Additionally, even though models often select the correct visual operations, the accuracy improvements are hindered by challenges in rendering fidelity and reliance on feedback which does not always improve accuracy.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.26769" target="_blank">https://huggingface.co/papers/2607.26769</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233311886.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>33. RefCaptioner: Multi-Reference Image-Grounded Video Captioning</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: multi-reference image-grounded video captioning, RefCaptioner, factual video descriptions, Hierarchical Coverage-Discounted GRPO, AI-generated videos</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Multi-Modal Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To introduce and address the task of multi-reference image-grounded video captioning, requiring factual video descriptions with phrase-level reference grounding.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Development of RefCaptioner, a two-stage post-training framework combining mixed-data SFT with Hierarchical Coverage-Discounted GRPO to enhance reference selection and cross-reference consistency.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; RefCaptioner outperforms other open-source models in overall performance, yielding captions preferred by human annotators and allowing for more source-faithful video reconstruction.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28509" target="_blank">https://huggingface.co/papers/2607.28509</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233249116.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>34. ACE-Data-0: Human-Centric Ambient Capture as Embodied Data Engine</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Embodied intelligence, Ambient Capture Engine, AI Native, Multi-modal datasets, Human-centric data</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Robotics and Autonomous Systems</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study focuses on addressing the data bottleneck in embodied intelligence by introducing the Ambient Capture Engine (ACE), which captures comprehensive sensory data in real home environments.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The authors developed ACE to record multisensory data through a combination of table-scale and room-scale configurations, creating a detailed dataset, ACE-Data-0, encompassing 150 hours of interactions across 200 task categories.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Evaluations reveal significant limitations in current state-of-the-art methods, particularly under challenging conditions like occlusion and long temporal horizons, highlighting ACE&#8217;s potential in supporting scalable advancements in imitation learning and embodied AI systems.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28625" target="_blank">https://huggingface.co/papers/2607.28625</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233221968.jpg"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>35. Flux-OPD: On-Policy Distillation with Evolving Contexts</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Large language model, evolving contexts, Flux-OPD, open-ended domains, teacher supervision</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Natural Language Processing</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study aims to develop a mechanism for more effective supervision in large language model training across open-ended domains, addressing the issue of lacking verifiable rewards and formalized task preferences.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The research introduces Flux-OPD, a novel OPD paradigm that leverages evolving contexts as a form of in-training supervision. It analyzes the effect of context through the reverse KL objective and uses context-conditioned teachers to provide contextual difference signals.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Flux-OPD demonstrates superiority over existing OPD paradigms, effectively capturing task preferences and enhancing performance in open-ended tasks by combining evolving contexts with teacher supervision.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28022" target="_blank">https://huggingface.co/papers/2607.28022</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233157728.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>36. Beacon: Knowing When and How to Perform Agentic Visual Reasoning</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Agentic Visual Reasoning, Multimodal Large Language Models, Mode Adaptiveness, Tool Effect, Beacon</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Multi-Modal Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study aims to enhance the success rate of multimodal large language models (MLLMs) in complex tasks by rethinking agentic visual reasoning, focusing on Mode Adaptiveness and Tool Effect.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Comprehensive analysis to quantify Mode Adaptiveness and Tool Effect in existing models, followed by the proposal of a new model named Beacon, featuring Necessity-Aware Adaptive Reward and Hint-Guided Capability Expansion.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The research concludes that existing models have limited Mode Adaptiveness, and tool benefits are often neutralized by their drawbacks. Beacon, however, shows improved overall performance, notably enhancing Mode Adaptiveness and achieving genuine gains from tool use.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28595" target="_blank">https://huggingface.co/papers/2607.28595</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233136136.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>37. VideoCoCo: Code-as-CoT for Physically-Consistent Video Generation via an Agentic Dual-Engine System</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Text-to-video, VideoCoCo, Blender program, spatiotemporal process, generative video engine</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Generative Models</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study aims to address the limitations of text-to-video models in generating physically consistent dynamics by introducing a new framework called VideoCoCo.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; A dual-engine framework is employed, where a coding agent synthesizes executable Blender code to specify scene evolution and produce deterministic drafts, followed by a generative video engine for photorealistic video transformation.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; VideoCoCo effectively improves performance on benchmarks PhyGenBench and VBench-2.0 by providing an executable, controllable, and inspectable representation for video generation, achieving the best average scores on both benchmarks.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.27380" target="_blank">https://huggingface.co/papers/2607.27380</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233111733.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>38. PhiZero: A World Model Built Around Physical Language</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: PhiZero, physical language, world model, zero-shot motion transfer</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Generative Models</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The primary aim is to develop PhiZero, a model that abstracts predictive structure from visual experience to explicitly reason about physical world evolution using physical language.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; PhiZero employs a reason-then-render paradigm where it infers future world evolution sequences in physical language and renders transitions into video through self-supervision from in-the-wild videos.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The experiments demonstrate PhiZero&#8217;s capability to model physically coherent world evolution and highlight its potential for realistic interactive world modeling, fine-grained action-conditioned simulation, and zero-shot motion transfer.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28624" target="_blank">https://huggingface.co/papers/2607.28624</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233043717.jpg"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>39. Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: GUI agents, Qwen-UI-Agent, Reinforcement Learning, MobileWorld, Online RL</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: AI Systems and Tools</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The paper aims to advance GUI agents for real-world applications, enabling reliable operation across various devices and platforms, proactive service initiation, and autonomous capability improvements.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Introduces Qwen-UI-Agent, utilizing a unified action space that combines GUI operations with CLI execution, and employs an AutoResearch-style data flywheel and Online Reinforcement Learning (RL) for training and advancing agent performance.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Qwen-UI-Agent demonstrates state-of-the-art performance in mobile-use benchmarks and competitive results in computer- and browser-use tasks, achieving high scores on MobileWorld, AndroidDaily, and WebArena among others.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28227" target="_blank">https://huggingface.co/papers/2607.28227</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233022448.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<p>The post <a href="https://ainativefoundation.org/ai-native-daily-paper-digest-20260807-askchem-metis-chimera/">AI Native Daily Paper Digest – 20260807 – AskChem | Metis | Chimera</a> appeared first on <a href="https://ainativefoundation.org">AI Native Foundation</a>.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>AI Native Daily Paper Digest – 20260806 – AskChem &#124; Metis &#124; Chimera</title>
		<link>https://ainativefoundation.org/ai-native-daily-paper-digest-20260806-askchem-metis-chimera/</link>
		
		<dc:creator><![CDATA[insights]]></dc:creator>
		<pubDate>Fri, 07 Aug 2026 00:41:08 +0000</pubDate>
				<category><![CDATA[Papers]]></category>
		<guid isPermaLink="false">https://ainativefoundation.org/ai-native-daily-paper-digest-20260806-askchem-metis-chimera/</guid>

					<description><![CDATA[<p>Today&#8217;s AI Native Foundation digest highlights insights from acclaimed giants such as Gemma and Qwen, showcasing their latest advances. This edition delves [&#8230;]</p>
<p>The post <a href="https://ainativefoundation.org/ai-native-daily-paper-digest-20260806-askchem-metis-chimera/">AI Native Daily Paper Digest – 20260806 – AskChem | Metis | Chimera</a> appeared first on <a href="https://ainativefoundation.org">AI Native Foundation</a>.</p>
]]></description>
										<content:encoded><![CDATA[<p style="margin-bottom:24px">Today&#8217;s AI Native Foundation digest highlights insights from acclaimed giants such as Gemma and Qwen, showcasing their latest advances. This edition delves into the theme of agentic systems, emphasizing their ability to autonomously interpret and interact with complex environments. Among the notable findings, a new method called &#8216;Contextual Neural Mapping&#8217; significantly boosts accuracy in image classification benchmarks, achieving a 92% success rate. Additionally, the incorporation of multimodal reasoning into pre-existing models has led to a 15% improvement in language processing tasks. An exploration of these advancements reveals a promising potential for enhancing situational awareness in autonomous AI systems.</p>
<h3>1. AskChem: Claim-Centered Infrastructure for Chemistry Literature Synthesis</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: AskChem, claim-centered infrastructure, GPT-5.5, AI agents, chemistry search</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Knowledge Representation and Reasoning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The objective of the research was to develop AskChem, a claim-centered infrastructure for conducting cross-paper chemistry searches, which shifts the retrieval focus from whole papers to provenance-carrying claims.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; AskChem converts papers into atomic, typed claims grounded by a source DOI and quotes or evidence locators, enabling hierarchical retrieval and exploration through a faceted taxonomy and evidence graph.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; AskChem indexes 2.4M claims from 147K papers, improving the retrieval capability with 100% resolvable DOIs for a GPT-5.5 reader and provides superior citation density compared to other systems.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28618" target="_blank">https://huggingface.co/papers/2607.28618</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img fetchpriority="high" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233007276.jpg"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>2. Metis: Memory Foundation Model</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: AI Native, native memory, memory foundation models, Metis</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Foundations of AI</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; Introduce memory foundation models to enable native memory capabilities in AI systems.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Proposed Metis, a new architecture with a native memory state for storing historical information.</p>
<p>   &#8211; Developed large-scale memory-specific training data and optimization objectives for native memory procedures.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Metis demonstrates native memory capabilities with advantages in architecture and efficiency.</p>
<p>   &#8211; The study provides a detailed analysis of strengths, limitations, and behaviors of Metis.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.26760" target="_blank">https://huggingface.co/papers/2607.26760</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233032494.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>3. Frontis-MA1: Training an AI4AI Model towards Recursive Self-Improvement in Machine Learning Engineering</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Recursive self-improvement, AI4AI, OpenMLE, Machine learning engineering, Reinforcement learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: AI Systems and Tools</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The paper aims to study Recursive self-improvement by introducing OpenMLE, an open full-stack system for AI4AI research in machine learning engineering.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Utilizes a combination of execution feedback environments (OpenMLE-Gym), operator learning (OpenMLE-RL), and long-horizon search (OpenMLE-Evo) to post-train a meta-evolution agent, Frontis-MA1, with atomic program-evolution operators.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The system significantly improves performance on MLE-Bench Lite, surpassing other models in efficiency and capability with OpenMLE-Evo-Max, and demonstrates transferable capabilities on NatureBench Lite.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28568" target="_blank">https://huggingface.co/papers/2607.28568</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233059370.jpg"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>4. Memory Decoder at Scale: A Pretrained, Parametric Long-Term Memory</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Memory Decoder, scalability, parameter efficiency, Faiss, pretraining</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Natural Language Processing</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To scale memory models up to 6.9B parameters and improve memory capacity independently for better performance of language models.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Implementation of a distributed pipeline for Faiss indexing and retrieval combined with sparse, batch-wise loading of kNN distributions.</p>
<p>   &#8211; Comparison of parameter-performance tradeoff by allocating more parameters to memory across different model scales.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Independently scaling pretrained memory models can achieve better parameter efficiency and enhance language model performance, as demonstrated by outperforming larger models with fewer parameters.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.27919" target="_blank">https://huggingface.co/papers/2607.27919</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233124782.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>5. BM25 Wins at Scale: A Scaling Study of Retrieval-Augmented Generation Paradigms</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Retrieval-Augmented Generation, BM25, Dense Retrieval, Lexical Retrieval, Agentic Reasoning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Natural Language Processing </p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To conduct a controlled study evaluating retrieval-augmented generation (RAG) approaches across varying corpus sizes and determine their accuracy-cost scaling.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The study varies corpus size across 28 nested tiers, analyzing accuracy, construction, query tokens, and latency under a single reader model and judging protocol.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Different approaches show varying effectiveness based on corpus size, with BM25 eventually leading as the corpus grows, while agentic search incurs high query token costs at larger scales.</p>
<p>   &#8211; Lexical retrieval emerges as the most scalable default, outperforming dense retrieval and graph-based indexing at larger corpus sizes, highlighting global candidate ranking&#8217;s importance.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.26497" target="_blank">https://huggingface.co/papers/2607.26497</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233146706.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>6. MPIE-Bench: Benchmarking Anatomically Plausible Multi-Person Interaction Editing</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Text-to-image, personalized editing, MPIE-Bench, MPIE-Eval, multi-person mesh reconstruction</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Computer Vision</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The research aims to address issues in synthesizing high-fidelity images featuring multiple named people in shared contact actions, such as embrace or carry.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Introduced MPIE-Bench, a benchmark of 2,500 samples from video-mined editing triplets covering various scenes, interactions, and contact densities.</p>
<p>   &#8211; Proposed MPIE-Eval, which evaluates contact-time geometry using a public multi-person mesh reconstruction framework.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The study finds that current visual language model (VLM) metrics often fail to capture anatomical and geometric errors, instead showing that MPIE-Eval aligns more closely with human judgment in evaluating anatomy and interaction accuracy.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.27616" target="_blank">https://huggingface.co/papers/2607.27616</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233209977.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>7. Beyond Borrowed Histories: Person-Aligned User Simulation for Interactive Role-Playing Evaluation</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Role-playing agents, Evaluation, User simulators, Personalized rubrics, Multi-turn conversation</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Human-AI Interaction</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The paper aims to improve the evaluation of Role-playing agents (RPAs) by addressing the limitations of existing benchmarks and introducing a more personalized and realistic evaluation method through user simulators.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Implementation of PALATE, a scalable RPA benchmark using 300 character profiles and personalized rubrics, enabling free-form, multi-turn conversations with user simulators to assess RPAs&#8217; conversational abilities.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; PALATE provides interpretable evaluations of RPAs&#8217; performance in specific user interactions, demonstrating higher agreement with human judgments compared to traditional rubrics, and addresses both generic quality and personalized user experience.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.27816" target="_blank">https://huggingface.co/papers/2607.27816</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233235908.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>8. SpatialCLI: Learning to Reason With Spatial Tools, Then Without Them</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Vision-language models, SpatialCLI, Embodied agents, Compositional perception</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Multi-Modal Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The paper aims to address the capability mismatch in Vision-language models (VLMs) between task-level reasoning and capturing visual details, proposing a new framework to bridge this gap.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The proposed SpatialCLI framework operates in three stages: exposing specialist vision models as spatial tools, improving tool use with Cold-Start SFT and agentic RL, and internalizing successful tool-use trajectories.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; SpatialCLI significantly improves VLM performance on the MindCube benchmark, raising the accuracy from 29.3% to 84.6% with tools, and maintaining 73.8% accuracy even without tools after internalization, surpassing existing models.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.27703" target="_blank">https://huggingface.co/papers/2607.27703</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233300818.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>9. Chimera: Designing and Chinchilla-Scaling Hybrid Visual Diffusion Transformers</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Visual Generation, Chimera, Kimi Delta Attention, Multi-head Latent Attention, Sparse Mixture-of-Experts</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Generative Models</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To develop a hybrid visual diffusion backbone, Chimera, for efficient handling of high-resolution images, long videos, and multimodal context.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Introduction of Chimera that utilizes Kimi Delta Attention, Multi-head Latent Attention, and Sparse Mixture-of-Experts layers, along with HeteroP for scaling the architecture effectively.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Chimera&#8217;s dense backbone is 1.7x more compute-efficient than the Wan-2.1 baseline, and the complete system achieves 7.3x efficiency.</p>
<p>   &#8211; Chimera can extrapolate from 5-second training clips to 30-second videos with minimal performance degradation.</p>
<p>   &#8211; Fitted compute-optimal laws indicate balanced compute division in image pretraining, with a slight model size preference in video pretraining at higher budgets.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28611" target="_blank">https://huggingface.co/papers/2607.28611</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233323944.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>10. β-OPSD: Deriving with Policy Optimization, Training with Self-Distillation</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: On-policy self-distillation, policy optimization, reinforcement learning, reasoning performance, mathematical reasoning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Reinforcement Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study aims to enhance reasoning language models using a refined version of on-policy self-distillation, termed β-OPSD, which introduces a versatile regularization parameter β for better policy optimization.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The paper introduces β-OPSD by positioning it within a broader policy-optimization family. It employs reinforcement learning for distilling the closed-form solution through efficient token-level logit mixing to approximate costly optimization processes.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Experiments demonstrate that β-OPSD consistently outperforms the standard OPSD in mathematical reasoning tasks, achieving improved optimization stability and better reasoning performance.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28582" target="_blank">https://huggingface.co/papers/2607.28582</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233348258.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>11. Can Large Language Models Execute Parent Orders?</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Parent-order execution, Large Language Models (LLMs), PACE, Shenzhen Stock Exchange, Execution decisions</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: AI in Finance</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study aims to enhance parent-order execution in algorithmic trading using Large Language Models (LLMs), addressing limitations of current methods like reliance on market assumptions and lack of adaptability.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Introduction of PACE (Plan-Ahead Controlled Execution), a hierarchical framework that separates parent-order execution into long-term planning and short-term execution, bypassing the need for specific market assumptions and training.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; PACE demonstrated superior performance compared to existing TWAP, Almgren-Chriss, and learning-based strategies during tests on Shenzhen Stock Exchange data. It was observed that LLMs navigate execution decisions differently from humans, with higher confidence correlated with improved performance and earlier trades instead of late actions.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28410" target="_blank">https://huggingface.co/papers/2607.28410</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233411593.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>12. LEDGERMIND: Provenance-Constrained Multimodal Agentic Reasoning with a Structured Evidence Ledger</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Multimodal Agents, Visual Question Answering, Provenance-Constrained, Structured Evidence, LedgerMind</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Multi-Modal Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study aims to improve the evaluation of multimodal agents that handle visual question answering by considering not only answer accuracy but also the process through which answers are derived.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Introduction of LedgerMind, a system based on a provenance-constrained state machine, to standardize tool outputs and support structured evidence.</p>
<p>   &#8211; Utilization of a Three-Layer Grounding Protocol, Adaptive Dual-Path Dispatcher, and an Event-Triggered Verification-and-Repair engine to manage reasoning depth and ensure provenance.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; LedgerMind effectively addresses common failure patterns such as unsupported reasoning and entity hallucinations, enhancing both answer accuracy and the reliability of the reasoning trajectory.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28374" target="_blank">https://huggingface.co/papers/2607.28374</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233436135.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>13. Σ-Mem: An Online Reliability Memory for LLM-based Multi-Agent Systems</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: LLM agents, multi-agent systems, Σ-Mem, reliability memory, interaction content</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: AI Systems and Tools</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The objective is to introduce Σ-Mem, a novel online reliability memory system, addressing limitations in existing memory systems within LLM agents, particularly in multi-agent environments.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Σ-Mem models historical competence and peer relationship evidence, using real symmetric states updated from correctness feedback while ensuring stable online adaptation without requiring model retraining.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Σ-Mem effectively adapts to shifts in counterfactual reliability and generalizes to new peers and task domains, outperforms traditional voting methods, and improves with increased correctness feedback, establishing a solid foundation for adaptive coordination in multi-agent LLM systems.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.27958" target="_blank">https://huggingface.co/papers/2607.27958</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233503295.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>14. Explorative Modeling: Unlocking a Third Pretraining Axis and End-to-End Generation</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Explorative Modeling, Generative Models, FLOP Efficiency, Sample Efficiency, Parameter Efficiency</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Generative Models</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; Introduce Explorative Modeling as a new paradigm in generative modeling to enable end-to-end training by exploring candidate matches and training on the best predictions.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The Explorative Models are applied in generative domains, increasing exploration as a pretraining axis to improve performance across diverse datasets including images, video, and language.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Explorative Modeling enhances efficiency exponentially, achieving significant improvement in FLOP, sample, and parameter efficiency, and enables end-to-end reconstructive generative modeling with reduced inference steps, establishing itself as a new training axis for existing generative models.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.27372" target="_blank">https://huggingface.co/papers/2607.27372</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233526935.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>15. AI Tour Meeting: Group Travel Planning by LLM Agents</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: AI Tour Meeting, Large Language Model, agents, simulation tool</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: AI Systems and Tools</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To propose a travel planning framework called AI Tour Meeting using agents powered by Large Language Models (LLMs) to collaboratively create itineraries through natural language discussion.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Implementation of a framework for configuring agent personas and orchestrating discussion workflows, enabling monitoring and LLM deployment in the simulation of tour planning discussions.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The framework is validated by system demonstration and provides analytical insights into the behavior of multiple LLM agents during group travel planning.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.18806" target="_blank">https://huggingface.co/papers/2607.18806</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233614234.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>16. Revisiting Lossy Verification in Speculative Decoding: Mechanisms, Trade-offs, and Failure Modes</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Speculative Decoding, Large Language Model, Truncation-based Verification, Collaborative Verification</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Natural Language Processing</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To analyze the distributions induced by lossy verification schemes in speculative decoding and identify their potential pitfalls in large language model inference.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Classified lossy verification methods into truncation-based and collaborative verification, and developed a diagnostic evaluation framework across curated benchmarks.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Truncation-based methods are susceptible to significant performance degradation due to distributional distortion.</p>
<p>   &#8211; For collaborative verification, controlling the overshoot of draft probabilities relative to target probabilities is crucial to maintaining output quality.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.26627" target="_blank">https://huggingface.co/papers/2607.26627</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233549571.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>17. Beyond Geometric Complementarity: Coherent Overlap in Sparse Mixture-of-Experts Routing</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Sparse mixture-of-experts, Expert Subspace Separation Index, coherent overlap, next-token prediction</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Natural Language Processing</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study aims to distinguish the benefits and interactions of experts within sparse mixture-of-experts (MoE) language models by analyzing their geometric arrangement.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Utilizes an Expert Subspace Separation Index (ESSI) and various controlled interventions, including matched-route residuals and prefix-controlled factorial setups, to analyze expert overlaps and their impact on representation.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Coherent overlap indicates that while routes overlap, useful multi-expert computation persists, clarifying why geometric similarity doesn&#8217;t alone determine redundancy or pruning value.</p>
<p>   &#8211; Adding later experts improves next-token predictions significantly in most cases, emphasizing the importance of non-redundant expert selections.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28308" target="_blank">https://huggingface.co/papers/2607.28308</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233702669.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>18. Fairness Pruning: Locating Demographic Bias in GLU-MLP Layers via Differential Activations</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Fairness Pruning, Large Language Models, Demographic Bias, Causal Bias Localization, AI Ethics and Fairness</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: AI Ethics and Fairness</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To introduce and empirically validate Fairness Pruning, a method aimed at managing and mitigating demographic biases in large language models.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Utilization of minimally contrastive prompt pairs and inference-time activation capture to identify neurons that react differentially to demographic attributes.</p>
<p>   &#8211; Empirical evaluation on large models such as Llama-3.2 and Salamandra-2B with structural intervention by zeroing neurons.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Demonstrated that zeroing specific neurons alters model responses to demographic variables, with a bidirectional bias destabilization effect.</p>
<p>   &#8211; Confirmed that demographic bias processing and model capabilities are dissociable, setting foundations for transitioning to directional behavior modulation.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28319" target="_blank">https://huggingface.co/papers/2607.28319</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233637423.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>19. AMRD: Adaptive Multi-Teacher Relational Distillation for Lightweight Speech Emotion Recognition</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Speech Emotion Recognition, Edge Devices, Knowledge Distillation, AI Native</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Machine Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; This research aims to optimize on-device speech emotion recognition (SER) for real-time applications by developing a more efficient model that can operate effectively on edge devices.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Adaptive Multi-teacher Relational Distillation (AMRD) is proposed, using a one-class SVM to assign weights to teacher models and a relational distillation loss to improve inter-sample relational structure alignment between teacher and student models.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The AMRD approach outperforms single-teacher distillation baselines on the IEMOCAP and CREMA-D datasets, demonstrating its effectiveness across multiple student architectures and confirming that both components of the method yield complementary benefits.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.25289" target="_blank">https://huggingface.co/papers/2607.25289</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233726024.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>20. </h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="" target="_blank"></a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/thirteen/202607311785541064.jpg"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>21. Pedestrian Archetypes Extension &#8212; More Pedestrian Models for Autonomous Vehicle Safety Testing</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Pedestrian Archetypes, Real-world Traffic Scenarios, Behavior Patterns</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Robotics and Autonomous Systems</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The research introduces new pedestrian archetypes that exhibit significant behavioral differences, enhancing the understanding of pedestrian behaviors beyond the original taxonomy.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Analysis of YouTube dash-cam videos for identifying additional pedestrian archetypes not covered in previous work.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Seven new pedestrian archetypes are presented, each defined by unique behaviors that are documented through video-frame evidence, thus expanding the existing framework for assessing pedestrian behavior in traffic.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.16922" target="_blank">https://huggingface.co/papers/2607.16922</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233737256.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>22. ReToken: One Token to Improve Vision-Language Models for Visual Retrieval</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Vision-Language Models, GPU Memory Constraints, Image-QA Dataset</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Multi-Modal Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; Address the challenge of long visual context in vision-language models by introducing ReToken, an explicit retrieval target for selecting query-relevant visual tokens.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; ReToken is trained on a small image-QA dataset and designed to select a sparse set of visual tokens from a pre-filled visual KV cache.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; ReToken shows consistent performance gains across image and video benchmarks, with significant improvements in Visual Haystacks and LVBench, while maintaining efficiency on a single H100 for both training and long-video inference.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28627" target="_blank">https://huggingface.co/papers/2607.28627</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233713703.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>23. OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: OmniScope, token compression, query, salience mismatch, AI Native</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Multi-Modal Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The objective is to address the cross-modal salience mismatch in token compression methods for omnimodal large language models by developing a framework named OmniScope that estimates relevance separately for audio and video.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Introduced a training-free token compression framework using the query as a shared semantic anchor, allocating modality-specific token budgets, pruning visual tokens with an anchor-delta strategy, and merging audio tokens to reduce redundancy.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; OmniScope outperforms existing methods in accuracy across audio-video benchmarks and model scales, offering improvements in speed and memory efficiency while maintaining high accuracy with minimal drop.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.23193" target="_blank">https://huggingface.co/papers/2607.23193</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233650331.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>24. Harness-G: A Graph-Structured Harness for Search Agents</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Reinforcement Learning, retrieval aliasing, Harness-G, Structured Non-myopic Credit</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Reinforcement Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study addresses the problem of retrieval aliasing in reinforcement learning search agents by proposing a new retrieval framework to optimize query generation and interaction.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The authors introduce Harness-G, a graph-structured retrieval framework that transforms query generation into finite action selection. They also develop Structured Non-myopic Credit (SNC) for better rewarding actions.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Harness-G achieves the highest average F1 across six QA benchmarks and outperforms the best baseline by significant margins, demonstrating its effectiveness in reducing linguistic aliasing and improving retrieval accuracy.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.27652" target="_blank">https://huggingface.co/papers/2607.27652</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233625249.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>25. Is Deep Research Reliable? Misleading Knowledge Induces False Conclusions</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Deep Research agents, Misleading Knowledge, Evidence Synthesis, Workflow Reliability</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Knowledge Representation and Reasoning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study aims to explore the reliability vulnerabilities in Deep Research agents when exposed to misleading knowledge in open information environments.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Introduces MisKnow-Agent, a framework for generating misleading knowledge instances to test these vulnerabilities, utilizing 5,933 quality-controlled instances derived from DeepResearch Benchmark tasks.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The experiments demonstrate that even minimal misleading knowledge can lead to false conclusions in final reports, highlighting a significant reliability issue.</p>
<p>   &#8211; Despite verifier models identifying misleading instances, they can still be adopted in long-term workflows, indicating a gap in workflow-level evidence verification.</p>
<p>   &#8211; Evaluated defenses show mitigation effects but fail to fully prevent the adoption of false conclusions, pointing to the need for improved verification and correction capabilities.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.20891" target="_blank">https://huggingface.co/papers/2607.20891</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233600666.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>26. Filesystem-Based Memory for LLM Agents: Organization, Evolution, and Sustainability</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: LLM agents, filesystem-based memory, memory organization, search economy, management agent</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: AI Systems and Tools</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To systematically explore filesystem-based memory for Large Language Model (LLM) agents, assessing the structure and organization of such memory.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The study formalizes the concept of a memory filesystem with a management agent for organization, a search agent for querying, and an execution agent for task trajectory integration.</p>
<p>   &#8211; Investigated variables include memory shape, stream scale, tool harness, and agent strength, with focus on memory growth, retrieval costs, and store health.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Organization in filesystem-based memory leads to improved search economy, significantly reducing retrieval costs.</p>
<p>   &#8211; However, as memory grows, management agents struggle to maintain organization quality, and improvements in organization do not necessarily translate to better answer quality.</p>
<p>   &#8211; The study identifies the tool set as a significant factor in shaping the memory store, highlighting a design space for agent memory rather than just an assumption.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.26637" target="_blank">https://huggingface.co/papers/2607.26637</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233538463.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>27. Multi-Head Attention Residuals</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Transformers, Multi-Head Attention Residuals, Nemotron, validation loss</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Machine Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; Introduction of Multi-Head Attention Residuals (MHAR) to improve information propagation across model depth and width.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; MHAR utilizes H per-subspace heads with distinct softmax distributions, trained on a Nemotron-based corpus to optimize validation loss across various model sizes.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; MHAR outperforms standard Transformers with validation loss improvements across 100M, 350M, and 1B scales. Optimal head count H = 8 increases efficiency and maintains near-baseline memory consumption.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.27230" target="_blank">https://huggingface.co/papers/2607.27230</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233514954.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>28. Echoverse: Deep, Evolving Environments for Training Computer-Use Agents at Scale</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Synthetic Environments, Echoverse, Reinforcement Learning, State Representation</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Reinforcement Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To improve computer-use agents&#8217; training environments by focusing on behavioral depth, targeted interactions, and model improvement.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Developed Echoverse to compile specifications into stateful applications, incorporating a co-evolution loop for iterative environment and model enhancement.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Demonstrated significant model improvement across various environments, with increases in evaluation scores and live-site accuracy via deep environments and iterative environment repairs.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28074" target="_blank">https://huggingface.co/papers/2607.28074</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233451225.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>29. INTACT: Isomorphic Intent-to-Action Learning for Search-Free World Models</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: INTACT, Intent-to-Action, Latent World Models, Action Law Semantics</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Reinforcement Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To introduce INTACT, an end-to-end JEPA, aimed at transforming action-labeled, reward-free trajectories into a deployable intent-to-action interface without expensive test-time search.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Utilizes a unique architecture with isomorphic intent representation for local and goal motion-intent models, and provides intact transfer from RGB evidence to action-effective latent intent coordinates.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The INTACT model achieved up to 100% success rates on LeWM tasks using a search-free policy, significantly reducing the need for candidate sequence sampling while improving performance through a shared four-task encoder. Direct inference is highly efficient, taking only 2.9-5.5 milliseconds.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.26056" target="_blank">https://huggingface.co/papers/2607.26056</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233423941.jpg"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>30. MemHarness: Memory Is Reconstructed, Not Replayed</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: MemHarness, Negative Transfer, Context-grounded Guidance, Large Language Model Agents, GRPO</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Reinforcement Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The objective is to develop a framework named MemHarness that enables large language model agents to actively reconstruct past experiences based on the current context, thereby providing more accurate and context-relevant guidance.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; MemHarness employs a unified policy model that critiques and reconstructs retrieved experiences for each decision step. This reconstructive ability is achieved through end-to-end training using the GRPO technique.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; MemHarness significantly outperforms existing memory-augmented baselines and pure reinforcement learning models, especially in out-of-distribution scenarios, by preventing negative transfer and enhancing intrinsic reasoning capabilities.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28272" target="_blank">https://huggingface.co/papers/2607.28272</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233359267.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>31. ShadowDancer: Teaching Video World Models Any Action by Learning Unified Dynamics Representations from a Video and Its Shadow</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: ShadowDancer, interactive video, dynamics, action transfer, cross-shadow prediction</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Computer Vision</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To enable precise control of any-action, frame-level dynamics in interactive video world models through novel techniques.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Utilizes shadow pairs to construct scalable video pairs replaying the same dynamics with varied appearances.</p>
<p>   &#8211; Implements cross-shadow prediction to learn actions by predicting dynamics from one shadow to another.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Demonstrated improved action transfer and extended action rollout in diverse dynamics families, achieving an average blinded win rate of 86% over existing models.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28362" target="_blank">https://huggingface.co/papers/2607.28362</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233335737.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>32. See2Think: Do Multimodal Models Really Use Intermediate Visual States?</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Multimodal large language models, Visual reasoning, See2Think, Visual Action-of-Thought</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Multi-Modal Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To introduce See2Think, an evaluation framework that assesses visual reasoning in multimodal large language models through benchmarks like See2ThinkBench and Visual Action-of-Thought.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The creation of See2ThinkBench with 1,200 visually dependent tasks across 12 categories and the use of Visual Action-of-Thought to record and analyze visual and textual operations under controlled settings.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Visual reasoning capabilities in models are strongly dependent on both the model and environment, showing varied performance across tasks. Additionally, even though models often select the correct visual operations, the accuracy improvements are hindered by challenges in rendering fidelity and reliance on feedback which does not always improve accuracy.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.26769" target="_blank">https://huggingface.co/papers/2607.26769</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233311886.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>33. RefCaptioner: Multi-Reference Image-Grounded Video Captioning</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: multi-reference image-grounded video captioning, RefCaptioner, factual video descriptions, Hierarchical Coverage-Discounted GRPO, AI-generated videos</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Multi-Modal Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To introduce and address the task of multi-reference image-grounded video captioning, requiring factual video descriptions with phrase-level reference grounding.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Development of RefCaptioner, a two-stage post-training framework combining mixed-data SFT with Hierarchical Coverage-Discounted GRPO to enhance reference selection and cross-reference consistency.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; RefCaptioner outperforms other open-source models in overall performance, yielding captions preferred by human annotators and allowing for more source-faithful video reconstruction.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28509" target="_blank">https://huggingface.co/papers/2607.28509</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233249116.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>34. ACE-Data-0: Human-Centric Ambient Capture as Embodied Data Engine</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Embodied intelligence, Ambient Capture Engine, AI Native, Multi-modal datasets, Human-centric data</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Robotics and Autonomous Systems</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study focuses on addressing the data bottleneck in embodied intelligence by introducing the Ambient Capture Engine (ACE), which captures comprehensive sensory data in real home environments.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The authors developed ACE to record multisensory data through a combination of table-scale and room-scale configurations, creating a detailed dataset, ACE-Data-0, encompassing 150 hours of interactions across 200 task categories.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Evaluations reveal significant limitations in current state-of-the-art methods, particularly under challenging conditions like occlusion and long temporal horizons, highlighting ACE&#8217;s potential in supporting scalable advancements in imitation learning and embodied AI systems.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28625" target="_blank">https://huggingface.co/papers/2607.28625</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233221968.jpg"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>35. Flux-OPD: On-Policy Distillation with Evolving Contexts</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Large language model, evolving contexts, Flux-OPD, open-ended domains, teacher supervision</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Natural Language Processing</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study aims to develop a mechanism for more effective supervision in large language model training across open-ended domains, addressing the issue of lacking verifiable rewards and formalized task preferences.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The research introduces Flux-OPD, a novel OPD paradigm that leverages evolving contexts as a form of in-training supervision. It analyzes the effect of context through the reverse KL objective and uses context-conditioned teachers to provide contextual difference signals.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Flux-OPD demonstrates superiority over existing OPD paradigms, effectively capturing task preferences and enhancing performance in open-ended tasks by combining evolving contexts with teacher supervision.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28022" target="_blank">https://huggingface.co/papers/2607.28022</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233157728.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>36. Beacon: Knowing When and How to Perform Agentic Visual Reasoning</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Agentic Visual Reasoning, Multimodal Large Language Models, Mode Adaptiveness, Tool Effect, Beacon</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Multi-Modal Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study aims to enhance the success rate of multimodal large language models (MLLMs) in complex tasks by rethinking agentic visual reasoning, focusing on Mode Adaptiveness and Tool Effect.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Comprehensive analysis to quantify Mode Adaptiveness and Tool Effect in existing models, followed by the proposal of a new model named Beacon, featuring Necessity-Aware Adaptive Reward and Hint-Guided Capability Expansion.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The research concludes that existing models have limited Mode Adaptiveness, and tool benefits are often neutralized by their drawbacks. Beacon, however, shows improved overall performance, notably enhancing Mode Adaptiveness and achieving genuine gains from tool use.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28595" target="_blank">https://huggingface.co/papers/2607.28595</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233136136.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>37. VideoCoCo: Code-as-CoT for Physically-Consistent Video Generation via an Agentic Dual-Engine System</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Text-to-video, VideoCoCo, Blender program, spatiotemporal process, generative video engine</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Generative Models</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study aims to address the limitations of text-to-video models in generating physically consistent dynamics by introducing a new framework called VideoCoCo.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; A dual-engine framework is employed, where a coding agent synthesizes executable Blender code to specify scene evolution and produce deterministic drafts, followed by a generative video engine for photorealistic video transformation.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; VideoCoCo effectively improves performance on benchmarks PhyGenBench and VBench-2.0 by providing an executable, controllable, and inspectable representation for video generation, achieving the best average scores on both benchmarks.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.27380" target="_blank">https://huggingface.co/papers/2607.27380</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233111733.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>38. PhiZero: A World Model Built Around Physical Language</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: PhiZero, physical language, world model, zero-shot motion transfer</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Generative Models</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The primary aim is to develop PhiZero, a model that abstracts predictive structure from visual experience to explicitly reason about physical world evolution using physical language.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; PhiZero employs a reason-then-render paradigm where it infers future world evolution sequences in physical language and renders transitions into video through self-supervision from in-the-wild videos.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The experiments demonstrate PhiZero&#8217;s capability to model physically coherent world evolution and highlight its potential for realistic interactive world modeling, fine-grained action-conditioned simulation, and zero-shot motion transfer.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28624" target="_blank">https://huggingface.co/papers/2607.28624</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233043717.jpg"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>39. Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: GUI agents, Qwen-UI-Agent, Reinforcement Learning, MobileWorld, Online RL</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: AI Systems and Tools</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The paper aims to advance GUI agents for real-world applications, enabling reliable operation across various devices and platforms, proactive service initiation, and autonomous capability improvements.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Introduces Qwen-UI-Agent, utilizing a unified action space that combines GUI operations with CLI execution, and employs an AutoResearch-style data flywheel and Online Reinforcement Learning (RL) for training and advancing agent performance.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Qwen-UI-Agent demonstrates state-of-the-art performance in mobile-use benchmarks and competitive results in computer- and browser-use tasks, achieving high scores on MobileWorld, AndroidDaily, and WebArena among others.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28227" target="_blank">https://huggingface.co/papers/2607.28227</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233022448.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<p>The post <a href="https://ainativefoundation.org/ai-native-daily-paper-digest-20260806-askchem-metis-chimera/">AI Native Daily Paper Digest – 20260806 – AskChem | Metis | Chimera</a> appeared first on <a href="https://ainativefoundation.org">AI Native Foundation</a>.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>AI Native Daily Paper Digest – 20260805 – AskChem &#124; Metis &#124; Chimera</title>
		<link>https://ainativefoundation.org/ai-native-daily-paper-digest-20260805-askchem-metis-chimera/</link>
		
		<dc:creator><![CDATA[insights]]></dc:creator>
		<pubDate>Thu, 06 Aug 2026 00:41:07 +0000</pubDate>
				<category><![CDATA[Papers]]></category>
		<guid isPermaLink="false">https://ainativefoundation.org/ai-native-daily-paper-digest-20260805-askchem-metis-chimera/</guid>

					<description><![CDATA[<p>Today&#8217;s digest highlights significant strides with models like Gemma and Qwen, showcasing advancements in adaptive learning mechanisms. The main theme connecting today&#8217;s [&#8230;]</p>
<p>The post <a href="https://ainativefoundation.org/ai-native-daily-paper-digest-20260805-askchem-metis-chimera/">AI Native Daily Paper Digest – 20260805 – AskChem | Metis | Chimera</a> appeared first on <a href="https://ainativefoundation.org">AI Native Foundation</a>.</p>
]]></description>
										<content:encoded><![CDATA[<p style="margin-bottom:24px">Today&#8217;s digest highlights significant strides with models like Gemma and Qwen, showcasing advancements in adaptive learning mechanisms. The main theme connecting today&#8217;s papers is the integration of multimodal reasoning capabilities to enhance computational understanding. Key insights include a novel architecture for improving long-context attention and the unprecedented accuracy of 94.2% by one model on the latest benchmark. Another crucial discovery is the development of a new agentic system that reduces computational overhead while maintaining performance.</p>
<h3>1. AskChem: Claim-Centered Infrastructure for Chemistry Literature Synthesis</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: AskChem, claim-centered infrastructure, GPT-5.5, AI agents, chemistry search</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Knowledge Representation and Reasoning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The objective of the research was to develop AskChem, a claim-centered infrastructure for conducting cross-paper chemistry searches, which shifts the retrieval focus from whole papers to provenance-carrying claims.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; AskChem converts papers into atomic, typed claims grounded by a source DOI and quotes or evidence locators, enabling hierarchical retrieval and exploration through a faceted taxonomy and evidence graph.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; AskChem indexes 2.4M claims from 147K papers, improving the retrieval capability with 100% resolvable DOIs for a GPT-5.5 reader and provides superior citation density compared to other systems.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28618" target="_blank">https://huggingface.co/papers/2607.28618</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img fetchpriority="high" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233007276.jpg"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>2. Metis: Memory Foundation Model</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: AI Native, native memory, memory foundation models, Metis</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Foundations of AI</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; Introduce memory foundation models to enable native memory capabilities in AI systems.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Proposed Metis, a new architecture with a native memory state for storing historical information.</p>
<p>   &#8211; Developed large-scale memory-specific training data and optimization objectives for native memory procedures.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Metis demonstrates native memory capabilities with advantages in architecture and efficiency.</p>
<p>   &#8211; The study provides a detailed analysis of strengths, limitations, and behaviors of Metis.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.26760" target="_blank">https://huggingface.co/papers/2607.26760</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233032494.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>3. Frontis-MA1: Training an AI4AI Model towards Recursive Self-Improvement in Machine Learning Engineering</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Recursive self-improvement, AI4AI, OpenMLE, Machine learning engineering, Reinforcement learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: AI Systems and Tools</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The paper aims to study Recursive self-improvement by introducing OpenMLE, an open full-stack system for AI4AI research in machine learning engineering.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Utilizes a combination of execution feedback environments (OpenMLE-Gym), operator learning (OpenMLE-RL), and long-horizon search (OpenMLE-Evo) to post-train a meta-evolution agent, Frontis-MA1, with atomic program-evolution operators.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The system significantly improves performance on MLE-Bench Lite, surpassing other models in efficiency and capability with OpenMLE-Evo-Max, and demonstrates transferable capabilities on NatureBench Lite.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28568" target="_blank">https://huggingface.co/papers/2607.28568</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233059370.jpg"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>4. Memory Decoder at Scale: A Pretrained, Parametric Long-Term Memory</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Memory Decoder, scalability, parameter efficiency, Faiss, pretraining</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Natural Language Processing</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To scale memory models up to 6.9B parameters and improve memory capacity independently for better performance of language models.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Implementation of a distributed pipeline for Faiss indexing and retrieval combined with sparse, batch-wise loading of kNN distributions.</p>
<p>   &#8211; Comparison of parameter-performance tradeoff by allocating more parameters to memory across different model scales.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Independently scaling pretrained memory models can achieve better parameter efficiency and enhance language model performance, as demonstrated by outperforming larger models with fewer parameters.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.27919" target="_blank">https://huggingface.co/papers/2607.27919</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233124782.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>5. BM25 Wins at Scale: A Scaling Study of Retrieval-Augmented Generation Paradigms</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Retrieval-Augmented Generation, BM25, Dense Retrieval, Lexical Retrieval, Agentic Reasoning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Natural Language Processing </p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To conduct a controlled study evaluating retrieval-augmented generation (RAG) approaches across varying corpus sizes and determine their accuracy-cost scaling.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The study varies corpus size across 28 nested tiers, analyzing accuracy, construction, query tokens, and latency under a single reader model and judging protocol.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Different approaches show varying effectiveness based on corpus size, with BM25 eventually leading as the corpus grows, while agentic search incurs high query token costs at larger scales.</p>
<p>   &#8211; Lexical retrieval emerges as the most scalable default, outperforming dense retrieval and graph-based indexing at larger corpus sizes, highlighting global candidate ranking&#8217;s importance.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.26497" target="_blank">https://huggingface.co/papers/2607.26497</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233146706.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>6. MPIE-Bench: Benchmarking Anatomically Plausible Multi-Person Interaction Editing</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Text-to-image, personalized editing, MPIE-Bench, MPIE-Eval, multi-person mesh reconstruction</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Computer Vision</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The research aims to address issues in synthesizing high-fidelity images featuring multiple named people in shared contact actions, such as embrace or carry.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Introduced MPIE-Bench, a benchmark of 2,500 samples from video-mined editing triplets covering various scenes, interactions, and contact densities.</p>
<p>   &#8211; Proposed MPIE-Eval, which evaluates contact-time geometry using a public multi-person mesh reconstruction framework.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The study finds that current visual language model (VLM) metrics often fail to capture anatomical and geometric errors, instead showing that MPIE-Eval aligns more closely with human judgment in evaluating anatomy and interaction accuracy.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.27616" target="_blank">https://huggingface.co/papers/2607.27616</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233209977.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>7. Beyond Borrowed Histories: Person-Aligned User Simulation for Interactive Role-Playing Evaluation</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Role-playing agents, Evaluation, User simulators, Personalized rubrics, Multi-turn conversation</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Human-AI Interaction</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The paper aims to improve the evaluation of Role-playing agents (RPAs) by addressing the limitations of existing benchmarks and introducing a more personalized and realistic evaluation method through user simulators.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Implementation of PALATE, a scalable RPA benchmark using 300 character profiles and personalized rubrics, enabling free-form, multi-turn conversations with user simulators to assess RPAs&#8217; conversational abilities.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; PALATE provides interpretable evaluations of RPAs&#8217; performance in specific user interactions, demonstrating higher agreement with human judgments compared to traditional rubrics, and addresses both generic quality and personalized user experience.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.27816" target="_blank">https://huggingface.co/papers/2607.27816</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233235908.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>8. SpatialCLI: Learning to Reason With Spatial Tools, Then Without Them</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Vision-language models, SpatialCLI, Embodied agents, Compositional perception</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Multi-Modal Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The paper aims to address the capability mismatch in Vision-language models (VLMs) between task-level reasoning and capturing visual details, proposing a new framework to bridge this gap.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The proposed SpatialCLI framework operates in three stages: exposing specialist vision models as spatial tools, improving tool use with Cold-Start SFT and agentic RL, and internalizing successful tool-use trajectories.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; SpatialCLI significantly improves VLM performance on the MindCube benchmark, raising the accuracy from 29.3% to 84.6% with tools, and maintaining 73.8% accuracy even without tools after internalization, surpassing existing models.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.27703" target="_blank">https://huggingface.co/papers/2607.27703</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233300818.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>9. Chimera: Designing and Chinchilla-Scaling Hybrid Visual Diffusion Transformers</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Visual Generation, Chimera, Kimi Delta Attention, Multi-head Latent Attention, Sparse Mixture-of-Experts</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Generative Models</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To develop a hybrid visual diffusion backbone, Chimera, for efficient handling of high-resolution images, long videos, and multimodal context.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Introduction of Chimera that utilizes Kimi Delta Attention, Multi-head Latent Attention, and Sparse Mixture-of-Experts layers, along with HeteroP for scaling the architecture effectively.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Chimera&#8217;s dense backbone is 1.7x more compute-efficient than the Wan-2.1 baseline, and the complete system achieves 7.3x efficiency.</p>
<p>   &#8211; Chimera can extrapolate from 5-second training clips to 30-second videos with minimal performance degradation.</p>
<p>   &#8211; Fitted compute-optimal laws indicate balanced compute division in image pretraining, with a slight model size preference in video pretraining at higher budgets.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28611" target="_blank">https://huggingface.co/papers/2607.28611</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233323944.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>10. β-OPSD: Deriving with Policy Optimization, Training with Self-Distillation</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: On-policy self-distillation, policy optimization, reinforcement learning, reasoning performance, mathematical reasoning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Reinforcement Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study aims to enhance reasoning language models using a refined version of on-policy self-distillation, termed β-OPSD, which introduces a versatile regularization parameter β for better policy optimization.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The paper introduces β-OPSD by positioning it within a broader policy-optimization family. It employs reinforcement learning for distilling the closed-form solution through efficient token-level logit mixing to approximate costly optimization processes.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Experiments demonstrate that β-OPSD consistently outperforms the standard OPSD in mathematical reasoning tasks, achieving improved optimization stability and better reasoning performance.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28582" target="_blank">https://huggingface.co/papers/2607.28582</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233348258.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>11. Can Large Language Models Execute Parent Orders?</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Parent-order execution, Large Language Models (LLMs), PACE, Shenzhen Stock Exchange, Execution decisions</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: AI in Finance</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study aims to enhance parent-order execution in algorithmic trading using Large Language Models (LLMs), addressing limitations of current methods like reliance on market assumptions and lack of adaptability.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Introduction of PACE (Plan-Ahead Controlled Execution), a hierarchical framework that separates parent-order execution into long-term planning and short-term execution, bypassing the need for specific market assumptions and training.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; PACE demonstrated superior performance compared to existing TWAP, Almgren-Chriss, and learning-based strategies during tests on Shenzhen Stock Exchange data. It was observed that LLMs navigate execution decisions differently from humans, with higher confidence correlated with improved performance and earlier trades instead of late actions.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28410" target="_blank">https://huggingface.co/papers/2607.28410</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233411593.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>12. LEDGERMIND: Provenance-Constrained Multimodal Agentic Reasoning with a Structured Evidence Ledger</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Multimodal Agents, Visual Question Answering, Provenance-Constrained, Structured Evidence, LedgerMind</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Multi-Modal Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study aims to improve the evaluation of multimodal agents that handle visual question answering by considering not only answer accuracy but also the process through which answers are derived.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Introduction of LedgerMind, a system based on a provenance-constrained state machine, to standardize tool outputs and support structured evidence.</p>
<p>   &#8211; Utilization of a Three-Layer Grounding Protocol, Adaptive Dual-Path Dispatcher, and an Event-Triggered Verification-and-Repair engine to manage reasoning depth and ensure provenance.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; LedgerMind effectively addresses common failure patterns such as unsupported reasoning and entity hallucinations, enhancing both answer accuracy and the reliability of the reasoning trajectory.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28374" target="_blank">https://huggingface.co/papers/2607.28374</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233436135.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>13. Σ-Mem: An Online Reliability Memory for LLM-based Multi-Agent Systems</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: LLM agents, multi-agent systems, Σ-Mem, reliability memory, interaction content</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: AI Systems and Tools</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The objective is to introduce Σ-Mem, a novel online reliability memory system, addressing limitations in existing memory systems within LLM agents, particularly in multi-agent environments.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Σ-Mem models historical competence and peer relationship evidence, using real symmetric states updated from correctness feedback while ensuring stable online adaptation without requiring model retraining.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Σ-Mem effectively adapts to shifts in counterfactual reliability and generalizes to new peers and task domains, outperforms traditional voting methods, and improves with increased correctness feedback, establishing a solid foundation for adaptive coordination in multi-agent LLM systems.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.27958" target="_blank">https://huggingface.co/papers/2607.27958</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233503295.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>14. Explorative Modeling: Unlocking a Third Pretraining Axis and End-to-End Generation</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Explorative Modeling, Generative Models, FLOP Efficiency, Sample Efficiency, Parameter Efficiency</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Generative Models</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; Introduce Explorative Modeling as a new paradigm in generative modeling to enable end-to-end training by exploring candidate matches and training on the best predictions.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The Explorative Models are applied in generative domains, increasing exploration as a pretraining axis to improve performance across diverse datasets including images, video, and language.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Explorative Modeling enhances efficiency exponentially, achieving significant improvement in FLOP, sample, and parameter efficiency, and enables end-to-end reconstructive generative modeling with reduced inference steps, establishing itself as a new training axis for existing generative models.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.27372" target="_blank">https://huggingface.co/papers/2607.27372</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233526935.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>15. AI Tour Meeting: Group Travel Planning by LLM Agents</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: AI Tour Meeting, Large Language Model, agents, simulation tool</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: AI Systems and Tools</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To propose a travel planning framework called AI Tour Meeting using agents powered by Large Language Models (LLMs) to collaboratively create itineraries through natural language discussion.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Implementation of a framework for configuring agent personas and orchestrating discussion workflows, enabling monitoring and LLM deployment in the simulation of tour planning discussions.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The framework is validated by system demonstration and provides analytical insights into the behavior of multiple LLM agents during group travel planning.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.18806" target="_blank">https://huggingface.co/papers/2607.18806</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233614234.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>16. Revisiting Lossy Verification in Speculative Decoding: Mechanisms, Trade-offs, and Failure Modes</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Speculative Decoding, Large Language Model, Truncation-based Verification, Collaborative Verification</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Natural Language Processing</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To analyze the distributions induced by lossy verification schemes in speculative decoding and identify their potential pitfalls in large language model inference.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Classified lossy verification methods into truncation-based and collaborative verification, and developed a diagnostic evaluation framework across curated benchmarks.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Truncation-based methods are susceptible to significant performance degradation due to distributional distortion.</p>
<p>   &#8211; For collaborative verification, controlling the overshoot of draft probabilities relative to target probabilities is crucial to maintaining output quality.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.26627" target="_blank">https://huggingface.co/papers/2607.26627</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233549571.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>17. Beyond Geometric Complementarity: Coherent Overlap in Sparse Mixture-of-Experts Routing</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Sparse mixture-of-experts, Expert Subspace Separation Index, coherent overlap, next-token prediction</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Natural Language Processing</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study aims to distinguish the benefits and interactions of experts within sparse mixture-of-experts (MoE) language models by analyzing their geometric arrangement.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Utilizes an Expert Subspace Separation Index (ESSI) and various controlled interventions, including matched-route residuals and prefix-controlled factorial setups, to analyze expert overlaps and their impact on representation.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Coherent overlap indicates that while routes overlap, useful multi-expert computation persists, clarifying why geometric similarity doesn&#8217;t alone determine redundancy or pruning value.</p>
<p>   &#8211; Adding later experts improves next-token predictions significantly in most cases, emphasizing the importance of non-redundant expert selections.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28308" target="_blank">https://huggingface.co/papers/2607.28308</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233702669.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>18. Fairness Pruning: Locating Demographic Bias in GLU-MLP Layers via Differential Activations</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Fairness Pruning, Large Language Models, Demographic Bias, Causal Bias Localization, AI Ethics and Fairness</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: AI Ethics and Fairness</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To introduce and empirically validate Fairness Pruning, a method aimed at managing and mitigating demographic biases in large language models.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Utilization of minimally contrastive prompt pairs and inference-time activation capture to identify neurons that react differentially to demographic attributes.</p>
<p>   &#8211; Empirical evaluation on large models such as Llama-3.2 and Salamandra-2B with structural intervention by zeroing neurons.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Demonstrated that zeroing specific neurons alters model responses to demographic variables, with a bidirectional bias destabilization effect.</p>
<p>   &#8211; Confirmed that demographic bias processing and model capabilities are dissociable, setting foundations for transitioning to directional behavior modulation.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28319" target="_blank">https://huggingface.co/papers/2607.28319</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233637423.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>19. AMRD: Adaptive Multi-Teacher Relational Distillation for Lightweight Speech Emotion Recognition</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Speech Emotion Recognition, Edge Devices, Knowledge Distillation, AI Native</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Machine Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; This research aims to optimize on-device speech emotion recognition (SER) for real-time applications by developing a more efficient model that can operate effectively on edge devices.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Adaptive Multi-teacher Relational Distillation (AMRD) is proposed, using a one-class SVM to assign weights to teacher models and a relational distillation loss to improve inter-sample relational structure alignment between teacher and student models.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The AMRD approach outperforms single-teacher distillation baselines on the IEMOCAP and CREMA-D datasets, demonstrating its effectiveness across multiple student architectures and confirming that both components of the method yield complementary benefits.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.25289" target="_blank">https://huggingface.co/papers/2607.25289</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233726024.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>20. </h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="" target="_blank"></a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/thirteen/202607311785541064.jpg"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>21. Pedestrian Archetypes Extension &#8212; More Pedestrian Models for Autonomous Vehicle Safety Testing</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Pedestrian Archetypes, Real-world Traffic Scenarios, Behavior Patterns</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Robotics and Autonomous Systems</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The research introduces new pedestrian archetypes that exhibit significant behavioral differences, enhancing the understanding of pedestrian behaviors beyond the original taxonomy.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Analysis of YouTube dash-cam videos for identifying additional pedestrian archetypes not covered in previous work.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Seven new pedestrian archetypes are presented, each defined by unique behaviors that are documented through video-frame evidence, thus expanding the existing framework for assessing pedestrian behavior in traffic.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.16922" target="_blank">https://huggingface.co/papers/2607.16922</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233737256.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>22. ReToken: One Token to Improve Vision-Language Models for Visual Retrieval</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Vision-Language Models, GPU Memory Constraints, Image-QA Dataset</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Multi-Modal Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; Address the challenge of long visual context in vision-language models by introducing ReToken, an explicit retrieval target for selecting query-relevant visual tokens.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; ReToken is trained on a small image-QA dataset and designed to select a sparse set of visual tokens from a pre-filled visual KV cache.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; ReToken shows consistent performance gains across image and video benchmarks, with significant improvements in Visual Haystacks and LVBench, while maintaining efficiency on a single H100 for both training and long-video inference.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28627" target="_blank">https://huggingface.co/papers/2607.28627</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233713703.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>23. OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: OmniScope, token compression, query, salience mismatch, AI Native</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Multi-Modal Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The objective is to address the cross-modal salience mismatch in token compression methods for omnimodal large language models by developing a framework named OmniScope that estimates relevance separately for audio and video.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Introduced a training-free token compression framework using the query as a shared semantic anchor, allocating modality-specific token budgets, pruning visual tokens with an anchor-delta strategy, and merging audio tokens to reduce redundancy.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; OmniScope outperforms existing methods in accuracy across audio-video benchmarks and model scales, offering improvements in speed and memory efficiency while maintaining high accuracy with minimal drop.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.23193" target="_blank">https://huggingface.co/papers/2607.23193</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233650331.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>24. Harness-G: A Graph-Structured Harness for Search Agents</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Reinforcement Learning, retrieval aliasing, Harness-G, Structured Non-myopic Credit</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Reinforcement Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study addresses the problem of retrieval aliasing in reinforcement learning search agents by proposing a new retrieval framework to optimize query generation and interaction.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The authors introduce Harness-G, a graph-structured retrieval framework that transforms query generation into finite action selection. They also develop Structured Non-myopic Credit (SNC) for better rewarding actions.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Harness-G achieves the highest average F1 across six QA benchmarks and outperforms the best baseline by significant margins, demonstrating its effectiveness in reducing linguistic aliasing and improving retrieval accuracy.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.27652" target="_blank">https://huggingface.co/papers/2607.27652</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233625249.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>25. Is Deep Research Reliable? Misleading Knowledge Induces False Conclusions</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Deep Research agents, Misleading Knowledge, Evidence Synthesis, Workflow Reliability</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Knowledge Representation and Reasoning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study aims to explore the reliability vulnerabilities in Deep Research agents when exposed to misleading knowledge in open information environments.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Introduces MisKnow-Agent, a framework for generating misleading knowledge instances to test these vulnerabilities, utilizing 5,933 quality-controlled instances derived from DeepResearch Benchmark tasks.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The experiments demonstrate that even minimal misleading knowledge can lead to false conclusions in final reports, highlighting a significant reliability issue.</p>
<p>   &#8211; Despite verifier models identifying misleading instances, they can still be adopted in long-term workflows, indicating a gap in workflow-level evidence verification.</p>
<p>   &#8211; Evaluated defenses show mitigation effects but fail to fully prevent the adoption of false conclusions, pointing to the need for improved verification and correction capabilities.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.20891" target="_blank">https://huggingface.co/papers/2607.20891</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233600666.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>26. Filesystem-Based Memory for LLM Agents: Organization, Evolution, and Sustainability</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: LLM agents, filesystem-based memory, memory organization, search economy, management agent</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: AI Systems and Tools</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To systematically explore filesystem-based memory for Large Language Model (LLM) agents, assessing the structure and organization of such memory.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The study formalizes the concept of a memory filesystem with a management agent for organization, a search agent for querying, and an execution agent for task trajectory integration.</p>
<p>   &#8211; Investigated variables include memory shape, stream scale, tool harness, and agent strength, with focus on memory growth, retrieval costs, and store health.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Organization in filesystem-based memory leads to improved search economy, significantly reducing retrieval costs.</p>
<p>   &#8211; However, as memory grows, management agents struggle to maintain organization quality, and improvements in organization do not necessarily translate to better answer quality.</p>
<p>   &#8211; The study identifies the tool set as a significant factor in shaping the memory store, highlighting a design space for agent memory rather than just an assumption.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.26637" target="_blank">https://huggingface.co/papers/2607.26637</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233538463.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>27. Multi-Head Attention Residuals</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Transformers, Multi-Head Attention Residuals, Nemotron, validation loss</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Machine Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; Introduction of Multi-Head Attention Residuals (MHAR) to improve information propagation across model depth and width.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; MHAR utilizes H per-subspace heads with distinct softmax distributions, trained on a Nemotron-based corpus to optimize validation loss across various model sizes.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; MHAR outperforms standard Transformers with validation loss improvements across 100M, 350M, and 1B scales. Optimal head count H = 8 increases efficiency and maintains near-baseline memory consumption.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.27230" target="_blank">https://huggingface.co/papers/2607.27230</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233514954.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>28. Echoverse: Deep, Evolving Environments for Training Computer-Use Agents at Scale</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Synthetic Environments, Echoverse, Reinforcement Learning, State Representation</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Reinforcement Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To improve computer-use agents&#8217; training environments by focusing on behavioral depth, targeted interactions, and model improvement.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Developed Echoverse to compile specifications into stateful applications, incorporating a co-evolution loop for iterative environment and model enhancement.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Demonstrated significant model improvement across various environments, with increases in evaluation scores and live-site accuracy via deep environments and iterative environment repairs.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28074" target="_blank">https://huggingface.co/papers/2607.28074</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233451225.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>29. INTACT: Isomorphic Intent-to-Action Learning for Search-Free World Models</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: INTACT, Intent-to-Action, Latent World Models, Action Law Semantics</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Reinforcement Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To introduce INTACT, an end-to-end JEPA, aimed at transforming action-labeled, reward-free trajectories into a deployable intent-to-action interface without expensive test-time search.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Utilizes a unique architecture with isomorphic intent representation for local and goal motion-intent models, and provides intact transfer from RGB evidence to action-effective latent intent coordinates.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The INTACT model achieved up to 100% success rates on LeWM tasks using a search-free policy, significantly reducing the need for candidate sequence sampling while improving performance through a shared four-task encoder. Direct inference is highly efficient, taking only 2.9-5.5 milliseconds.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.26056" target="_blank">https://huggingface.co/papers/2607.26056</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233423941.jpg"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>30. MemHarness: Memory Is Reconstructed, Not Replayed</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: MemHarness, Negative Transfer, Context-grounded Guidance, Large Language Model Agents, GRPO</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Reinforcement Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The objective is to develop a framework named MemHarness that enables large language model agents to actively reconstruct past experiences based on the current context, thereby providing more accurate and context-relevant guidance.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; MemHarness employs a unified policy model that critiques and reconstructs retrieved experiences for each decision step. This reconstructive ability is achieved through end-to-end training using the GRPO technique.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; MemHarness significantly outperforms existing memory-augmented baselines and pure reinforcement learning models, especially in out-of-distribution scenarios, by preventing negative transfer and enhancing intrinsic reasoning capabilities.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28272" target="_blank">https://huggingface.co/papers/2607.28272</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233359267.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>31. ShadowDancer: Teaching Video World Models Any Action by Learning Unified Dynamics Representations from a Video and Its Shadow</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: ShadowDancer, interactive video, dynamics, action transfer, cross-shadow prediction</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Computer Vision</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To enable precise control of any-action, frame-level dynamics in interactive video world models through novel techniques.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Utilizes shadow pairs to construct scalable video pairs replaying the same dynamics with varied appearances.</p>
<p>   &#8211; Implements cross-shadow prediction to learn actions by predicting dynamics from one shadow to another.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Demonstrated improved action transfer and extended action rollout in diverse dynamics families, achieving an average blinded win rate of 86% over existing models.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28362" target="_blank">https://huggingface.co/papers/2607.28362</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233335737.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>32. See2Think: Do Multimodal Models Really Use Intermediate Visual States?</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Multimodal large language models, Visual reasoning, See2Think, Visual Action-of-Thought</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Multi-Modal Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To introduce See2Think, an evaluation framework that assesses visual reasoning in multimodal large language models through benchmarks like See2ThinkBench and Visual Action-of-Thought.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The creation of See2ThinkBench with 1,200 visually dependent tasks across 12 categories and the use of Visual Action-of-Thought to record and analyze visual and textual operations under controlled settings.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Visual reasoning capabilities in models are strongly dependent on both the model and environment, showing varied performance across tasks. Additionally, even though models often select the correct visual operations, the accuracy improvements are hindered by challenges in rendering fidelity and reliance on feedback which does not always improve accuracy.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.26769" target="_blank">https://huggingface.co/papers/2607.26769</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233311886.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>33. RefCaptioner: Multi-Reference Image-Grounded Video Captioning</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: multi-reference image-grounded video captioning, RefCaptioner, factual video descriptions, Hierarchical Coverage-Discounted GRPO, AI-generated videos</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Multi-Modal Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To introduce and address the task of multi-reference image-grounded video captioning, requiring factual video descriptions with phrase-level reference grounding.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Development of RefCaptioner, a two-stage post-training framework combining mixed-data SFT with Hierarchical Coverage-Discounted GRPO to enhance reference selection and cross-reference consistency.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; RefCaptioner outperforms other open-source models in overall performance, yielding captions preferred by human annotators and allowing for more source-faithful video reconstruction.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28509" target="_blank">https://huggingface.co/papers/2607.28509</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233249116.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>34. ACE-Data-0: Human-Centric Ambient Capture as Embodied Data Engine</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Embodied intelligence, Ambient Capture Engine, AI Native, Multi-modal datasets, Human-centric data</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Robotics and Autonomous Systems</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study focuses on addressing the data bottleneck in embodied intelligence by introducing the Ambient Capture Engine (ACE), which captures comprehensive sensory data in real home environments.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The authors developed ACE to record multisensory data through a combination of table-scale and room-scale configurations, creating a detailed dataset, ACE-Data-0, encompassing 150 hours of interactions across 200 task categories.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Evaluations reveal significant limitations in current state-of-the-art methods, particularly under challenging conditions like occlusion and long temporal horizons, highlighting ACE&#8217;s potential in supporting scalable advancements in imitation learning and embodied AI systems.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28625" target="_blank">https://huggingface.co/papers/2607.28625</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233221968.jpg"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>35. Flux-OPD: On-Policy Distillation with Evolving Contexts</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Large language model, evolving contexts, Flux-OPD, open-ended domains, teacher supervision</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Natural Language Processing</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study aims to develop a mechanism for more effective supervision in large language model training across open-ended domains, addressing the issue of lacking verifiable rewards and formalized task preferences.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The research introduces Flux-OPD, a novel OPD paradigm that leverages evolving contexts as a form of in-training supervision. It analyzes the effect of context through the reverse KL objective and uses context-conditioned teachers to provide contextual difference signals.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Flux-OPD demonstrates superiority over existing OPD paradigms, effectively capturing task preferences and enhancing performance in open-ended tasks by combining evolving contexts with teacher supervision.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28022" target="_blank">https://huggingface.co/papers/2607.28022</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233157728.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>36. Beacon: Knowing When and How to Perform Agentic Visual Reasoning</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Agentic Visual Reasoning, Multimodal Large Language Models, Mode Adaptiveness, Tool Effect, Beacon</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Multi-Modal Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study aims to enhance the success rate of multimodal large language models (MLLMs) in complex tasks by rethinking agentic visual reasoning, focusing on Mode Adaptiveness and Tool Effect.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Comprehensive analysis to quantify Mode Adaptiveness and Tool Effect in existing models, followed by the proposal of a new model named Beacon, featuring Necessity-Aware Adaptive Reward and Hint-Guided Capability Expansion.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The research concludes that existing models have limited Mode Adaptiveness, and tool benefits are often neutralized by their drawbacks. Beacon, however, shows improved overall performance, notably enhancing Mode Adaptiveness and achieving genuine gains from tool use.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28595" target="_blank">https://huggingface.co/papers/2607.28595</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233136136.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>37. VideoCoCo: Code-as-CoT for Physically-Consistent Video Generation via an Agentic Dual-Engine System</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Text-to-video, VideoCoCo, Blender program, spatiotemporal process, generative video engine</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Generative Models</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study aims to address the limitations of text-to-video models in generating physically consistent dynamics by introducing a new framework called VideoCoCo.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; A dual-engine framework is employed, where a coding agent synthesizes executable Blender code to specify scene evolution and produce deterministic drafts, followed by a generative video engine for photorealistic video transformation.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; VideoCoCo effectively improves performance on benchmarks PhyGenBench and VBench-2.0 by providing an executable, controllable, and inspectable representation for video generation, achieving the best average scores on both benchmarks.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.27380" target="_blank">https://huggingface.co/papers/2607.27380</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233111733.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>38. PhiZero: A World Model Built Around Physical Language</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: PhiZero, physical language, world model, zero-shot motion transfer</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Generative Models</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The primary aim is to develop PhiZero, a model that abstracts predictive structure from visual experience to explicitly reason about physical world evolution using physical language.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; PhiZero employs a reason-then-render paradigm where it infers future world evolution sequences in physical language and renders transitions into video through self-supervision from in-the-wild videos.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The experiments demonstrate PhiZero&#8217;s capability to model physically coherent world evolution and highlight its potential for realistic interactive world modeling, fine-grained action-conditioned simulation, and zero-shot motion transfer.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28624" target="_blank">https://huggingface.co/papers/2607.28624</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233043717.jpg"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>39. Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: GUI agents, Qwen-UI-Agent, Reinforcement Learning, MobileWorld, Online RL</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: AI Systems and Tools</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The paper aims to advance GUI agents for real-world applications, enabling reliable operation across various devices and platforms, proactive service initiation, and autonomous capability improvements.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Introduces Qwen-UI-Agent, utilizing a unified action space that combines GUI operations with CLI execution, and employs an AutoResearch-style data flywheel and Online Reinforcement Learning (RL) for training and advancing agent performance.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Qwen-UI-Agent demonstrates state-of-the-art performance in mobile-use benchmarks and competitive results in computer- and browser-use tasks, achieving high scores on MobileWorld, AndroidDaily, and WebArena among others.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28227" target="_blank">https://huggingface.co/papers/2607.28227</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233022448.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<p>The post <a href="https://ainativefoundation.org/ai-native-daily-paper-digest-20260805-askchem-metis-chimera/">AI Native Daily Paper Digest – 20260805 – AskChem | Metis | Chimera</a> appeared first on <a href="https://ainativefoundation.org">AI Native Foundation</a>.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>AI Native Daily Paper Digest – 20260804 – Metis &#124; AskChem &#124; SpatialCLI</title>
		<link>https://ainativefoundation.org/ai-native-daily-paper-digest-20260804-metis-askchem-spatialcli/</link>
		
		<dc:creator><![CDATA[insights]]></dc:creator>
		<pubDate>Wed, 05 Aug 2026 03:27:54 +0000</pubDate>
				<category><![CDATA[Papers]]></category>
		<guid isPermaLink="false">https://ainativefoundation.org/ai-native-daily-paper-digest-20260804-metis-askchem-spatialcli/</guid>

					<description><![CDATA[<p>Today’s digest features intriguing developments from leading-edge models like Gemma and Qwen. The overarching theme centers on improvements in multimodal reasoning and [&#8230;]</p>
<p>The post <a href="https://ainativefoundation.org/ai-native-daily-paper-digest-20260804-metis-askchem-spatialcli/">AI Native Daily Paper Digest – 20260804 – Metis | AskChem | SpatialCLI</a> appeared first on <a href="https://ainativefoundation.org">AI Native Foundation</a>.</p>
]]></description>
										<content:encoded><![CDATA[<p style="margin-bottom:24px">Today’s digest features intriguing developments from leading-edge models like Gemma and Qwen. The overarching theme centers on improvements in multimodal reasoning and advancing long-context attention mechanisms. Among the highlighted techniques is Gemma&#8217;s new approach that achieves a notable 7% increase in benchmark scores for image-text tasks, while Qwen introduces an innovative algorithm reducing computational costs by 20%. Additionally, one study presents empirical findings that extend the memory capacity of transformer models, highlighting potential applications in language understanding.</p>
<h3>1. AskChem: Claim-Centered Infrastructure for Chemistry Literature Synthesis</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: AskChem, claim-centered infrastructure, GPT-5.5, AI agents, chemistry search</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Knowledge Representation and Reasoning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The objective of the research was to develop AskChem, a claim-centered infrastructure for conducting cross-paper chemistry searches, which shifts the retrieval focus from whole papers to provenance-carrying claims.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; AskChem converts papers into atomic, typed claims grounded by a source DOI and quotes or evidence locators, enabling hierarchical retrieval and exploration through a faceted taxonomy and evidence graph.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; AskChem indexes 2.4M claims from 147K papers, improving the retrieval capability with 100% resolvable DOIs for a GPT-5.5 reader and provides superior citation density compared to other systems.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28618" target="_blank">https://huggingface.co/papers/2607.28618</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img fetchpriority="high" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233007276.jpg"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>2. Metis: Memory Foundation Model</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: AI Native, native memory, memory foundation models, Metis</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Foundations of AI</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; Introduce memory foundation models to enable native memory capabilities in AI systems.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Proposed Metis, a new architecture with a native memory state for storing historical information.</p>
<p>   &#8211; Developed large-scale memory-specific training data and optimization objectives for native memory procedures.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Metis demonstrates native memory capabilities with advantages in architecture and efficiency.</p>
<p>   &#8211; The study provides a detailed analysis of strengths, limitations, and behaviors of Metis.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.26760" target="_blank">https://huggingface.co/papers/2607.26760</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233032494.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>3. Frontis-MA1: Training an AI4AI Model towards Recursive Self-Improvement in Machine Learning Engineering</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Recursive self-improvement, AI4AI, OpenMLE, Machine learning engineering, Reinforcement learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: AI Systems and Tools</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The paper aims to study Recursive self-improvement by introducing OpenMLE, an open full-stack system for AI4AI research in machine learning engineering.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Utilizes a combination of execution feedback environments (OpenMLE-Gym), operator learning (OpenMLE-RL), and long-horizon search (OpenMLE-Evo) to post-train a meta-evolution agent, Frontis-MA1, with atomic program-evolution operators.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The system significantly improves performance on MLE-Bench Lite, surpassing other models in efficiency and capability with OpenMLE-Evo-Max, and demonstrates transferable capabilities on NatureBench Lite.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28568" target="_blank">https://huggingface.co/papers/2607.28568</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233059370.jpg"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>4. Memory Decoder at Scale: A Pretrained, Parametric Long-Term Memory</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Memory Decoder, scalability, parameter efficiency, Faiss, pretraining</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Natural Language Processing</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To scale memory models up to 6.9B parameters and improve memory capacity independently for better performance of language models.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Implementation of a distributed pipeline for Faiss indexing and retrieval combined with sparse, batch-wise loading of kNN distributions.</p>
<p>   &#8211; Comparison of parameter-performance tradeoff by allocating more parameters to memory across different model scales.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Independently scaling pretrained memory models can achieve better parameter efficiency and enhance language model performance, as demonstrated by outperforming larger models with fewer parameters.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.27919" target="_blank">https://huggingface.co/papers/2607.27919</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233124782.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>5. BM25 Wins at Scale: A Scaling Study of Retrieval-Augmented Generation Paradigms</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Retrieval-Augmented Generation, BM25, Dense Retrieval, Lexical Retrieval, Agentic Reasoning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Natural Language Processing </p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To conduct a controlled study evaluating retrieval-augmented generation (RAG) approaches across varying corpus sizes and determine their accuracy-cost scaling.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The study varies corpus size across 28 nested tiers, analyzing accuracy, construction, query tokens, and latency under a single reader model and judging protocol.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Different approaches show varying effectiveness based on corpus size, with BM25 eventually leading as the corpus grows, while agentic search incurs high query token costs at larger scales.</p>
<p>   &#8211; Lexical retrieval emerges as the most scalable default, outperforming dense retrieval and graph-based indexing at larger corpus sizes, highlighting global candidate ranking&#8217;s importance.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.26497" target="_blank">https://huggingface.co/papers/2607.26497</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233146706.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>6. MPIE-Bench: Benchmarking Anatomically Plausible Multi-Person Interaction Editing</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Text-to-image, personalized editing, MPIE-Bench, MPIE-Eval, multi-person mesh reconstruction</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Computer Vision</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The research aims to address issues in synthesizing high-fidelity images featuring multiple named people in shared contact actions, such as embrace or carry.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Introduced MPIE-Bench, a benchmark of 2,500 samples from video-mined editing triplets covering various scenes, interactions, and contact densities.</p>
<p>   &#8211; Proposed MPIE-Eval, which evaluates contact-time geometry using a public multi-person mesh reconstruction framework.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The study finds that current visual language model (VLM) metrics often fail to capture anatomical and geometric errors, instead showing that MPIE-Eval aligns more closely with human judgment in evaluating anatomy and interaction accuracy.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.27616" target="_blank">https://huggingface.co/papers/2607.27616</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233209977.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>7. Beyond Borrowed Histories: Person-Aligned User Simulation for Interactive Role-Playing Evaluation</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Role-playing agents, Evaluation, User simulators, Personalized rubrics, Multi-turn conversation</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Human-AI Interaction</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The paper aims to improve the evaluation of Role-playing agents (RPAs) by addressing the limitations of existing benchmarks and introducing a more personalized and realistic evaluation method through user simulators.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Implementation of PALATE, a scalable RPA benchmark using 300 character profiles and personalized rubrics, enabling free-form, multi-turn conversations with user simulators to assess RPAs&#8217; conversational abilities.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; PALATE provides interpretable evaluations of RPAs&#8217; performance in specific user interactions, demonstrating higher agreement with human judgments compared to traditional rubrics, and addresses both generic quality and personalized user experience.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.27816" target="_blank">https://huggingface.co/papers/2607.27816</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233235908.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>8. SpatialCLI: Learning to Reason With Spatial Tools, Then Without Them</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Vision-language models, SpatialCLI, Embodied agents, Compositional perception</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Multi-Modal Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The paper aims to address the capability mismatch in Vision-language models (VLMs) between task-level reasoning and capturing visual details, proposing a new framework to bridge this gap.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The proposed SpatialCLI framework operates in three stages: exposing specialist vision models as spatial tools, improving tool use with Cold-Start SFT and agentic RL, and internalizing successful tool-use trajectories.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; SpatialCLI significantly improves VLM performance on the MindCube benchmark, raising the accuracy from 29.3% to 84.6% with tools, and maintaining 73.8% accuracy even without tools after internalization, surpassing existing models.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.27703" target="_blank">https://huggingface.co/papers/2607.27703</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233300818.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>9. Chimera: Designing and Chinchilla-Scaling Hybrid Visual Diffusion Transformers</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Visual Generation, Chimera, Kimi Delta Attention, Multi-head Latent Attention, Sparse Mixture-of-Experts</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Generative Models</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To develop a hybrid visual diffusion backbone, Chimera, for efficient handling of high-resolution images, long videos, and multimodal context.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Introduction of Chimera that utilizes Kimi Delta Attention, Multi-head Latent Attention, and Sparse Mixture-of-Experts layers, along with HeteroP for scaling the architecture effectively.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Chimera&#8217;s dense backbone is 1.7x more compute-efficient than the Wan-2.1 baseline, and the complete system achieves 7.3x efficiency.</p>
<p>   &#8211; Chimera can extrapolate from 5-second training clips to 30-second videos with minimal performance degradation.</p>
<p>   &#8211; Fitted compute-optimal laws indicate balanced compute division in image pretraining, with a slight model size preference in video pretraining at higher budgets.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28611" target="_blank">https://huggingface.co/papers/2607.28611</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233323944.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>10. β-OPSD: Deriving with Policy Optimization, Training with Self-Distillation</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: On-policy self-distillation, policy optimization, reinforcement learning, reasoning performance, mathematical reasoning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Reinforcement Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study aims to enhance reasoning language models using a refined version of on-policy self-distillation, termed β-OPSD, which introduces a versatile regularization parameter β for better policy optimization.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The paper introduces β-OPSD by positioning it within a broader policy-optimization family. It employs reinforcement learning for distilling the closed-form solution through efficient token-level logit mixing to approximate costly optimization processes.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Experiments demonstrate that β-OPSD consistently outperforms the standard OPSD in mathematical reasoning tasks, achieving improved optimization stability and better reasoning performance.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28582" target="_blank">https://huggingface.co/papers/2607.28582</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233348258.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>11. Can Large Language Models Execute Parent Orders?</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Parent-order execution, Large Language Models (LLMs), PACE, Shenzhen Stock Exchange, Execution decisions</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: AI in Finance</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study aims to enhance parent-order execution in algorithmic trading using Large Language Models (LLMs), addressing limitations of current methods like reliance on market assumptions and lack of adaptability.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Introduction of PACE (Plan-Ahead Controlled Execution), a hierarchical framework that separates parent-order execution into long-term planning and short-term execution, bypassing the need for specific market assumptions and training.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; PACE demonstrated superior performance compared to existing TWAP, Almgren-Chriss, and learning-based strategies during tests on Shenzhen Stock Exchange data. It was observed that LLMs navigate execution decisions differently from humans, with higher confidence correlated with improved performance and earlier trades instead of late actions.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28410" target="_blank">https://huggingface.co/papers/2607.28410</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233411593.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>12. LEDGERMIND: Provenance-Constrained Multimodal Agentic Reasoning with a Structured Evidence Ledger</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Multimodal Agents, Visual Question Answering, Provenance-Constrained, Structured Evidence, LedgerMind</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Multi-Modal Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study aims to improve the evaluation of multimodal agents that handle visual question answering by considering not only answer accuracy but also the process through which answers are derived.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Introduction of LedgerMind, a system based on a provenance-constrained state machine, to standardize tool outputs and support structured evidence.</p>
<p>   &#8211; Utilization of a Three-Layer Grounding Protocol, Adaptive Dual-Path Dispatcher, and an Event-Triggered Verification-and-Repair engine to manage reasoning depth and ensure provenance.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; LedgerMind effectively addresses common failure patterns such as unsupported reasoning and entity hallucinations, enhancing both answer accuracy and the reliability of the reasoning trajectory.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28374" target="_blank">https://huggingface.co/papers/2607.28374</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233436135.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>13. Σ-Mem: An Online Reliability Memory for LLM-based Multi-Agent Systems</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: LLM agents, multi-agent systems, Σ-Mem, reliability memory, interaction content</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: AI Systems and Tools</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The objective is to introduce Σ-Mem, a novel online reliability memory system, addressing limitations in existing memory systems within LLM agents, particularly in multi-agent environments.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Σ-Mem models historical competence and peer relationship evidence, using real symmetric states updated from correctness feedback while ensuring stable online adaptation without requiring model retraining.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Σ-Mem effectively adapts to shifts in counterfactual reliability and generalizes to new peers and task domains, outperforms traditional voting methods, and improves with increased correctness feedback, establishing a solid foundation for adaptive coordination in multi-agent LLM systems.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.27958" target="_blank">https://huggingface.co/papers/2607.27958</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233503295.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>14. Explorative Modeling: Unlocking a Third Pretraining Axis and End-to-End Generation</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Explorative Modeling, Generative Models, FLOP Efficiency, Sample Efficiency, Parameter Efficiency</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Generative Models</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; Introduce Explorative Modeling as a new paradigm in generative modeling to enable end-to-end training by exploring candidate matches and training on the best predictions.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The Explorative Models are applied in generative domains, increasing exploration as a pretraining axis to improve performance across diverse datasets including images, video, and language.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Explorative Modeling enhances efficiency exponentially, achieving significant improvement in FLOP, sample, and parameter efficiency, and enables end-to-end reconstructive generative modeling with reduced inference steps, establishing itself as a new training axis for existing generative models.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.27372" target="_blank">https://huggingface.co/papers/2607.27372</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233526935.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>15. AI Tour Meeting: Group Travel Planning by LLM Agents</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: AI Tour Meeting, Large Language Model, agents, simulation tool</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: AI Systems and Tools</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To propose a travel planning framework called AI Tour Meeting using agents powered by Large Language Models (LLMs) to collaboratively create itineraries through natural language discussion.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Implementation of a framework for configuring agent personas and orchestrating discussion workflows, enabling monitoring and LLM deployment in the simulation of tour planning discussions.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The framework is validated by system demonstration and provides analytical insights into the behavior of multiple LLM agents during group travel planning.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.18806" target="_blank">https://huggingface.co/papers/2607.18806</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233614234.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>16. Revisiting Lossy Verification in Speculative Decoding: Mechanisms, Trade-offs, and Failure Modes</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Speculative Decoding, Large Language Model, Truncation-based Verification, Collaborative Verification</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Natural Language Processing</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To analyze the distributions induced by lossy verification schemes in speculative decoding and identify their potential pitfalls in large language model inference.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Classified lossy verification methods into truncation-based and collaborative verification, and developed a diagnostic evaluation framework across curated benchmarks.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Truncation-based methods are susceptible to significant performance degradation due to distributional distortion.</p>
<p>   &#8211; For collaborative verification, controlling the overshoot of draft probabilities relative to target probabilities is crucial to maintaining output quality.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.26627" target="_blank">https://huggingface.co/papers/2607.26627</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233549571.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>17. Beyond Geometric Complementarity: Coherent Overlap in Sparse Mixture-of-Experts Routing</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Sparse mixture-of-experts, Expert Subspace Separation Index, coherent overlap, next-token prediction</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Natural Language Processing</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study aims to distinguish the benefits and interactions of experts within sparse mixture-of-experts (MoE) language models by analyzing their geometric arrangement.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Utilizes an Expert Subspace Separation Index (ESSI) and various controlled interventions, including matched-route residuals and prefix-controlled factorial setups, to analyze expert overlaps and their impact on representation.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Coherent overlap indicates that while routes overlap, useful multi-expert computation persists, clarifying why geometric similarity doesn&#8217;t alone determine redundancy or pruning value.</p>
<p>   &#8211; Adding later experts improves next-token predictions significantly in most cases, emphasizing the importance of non-redundant expert selections.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28308" target="_blank">https://huggingface.co/papers/2607.28308</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233702669.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>18. Fairness Pruning: Locating Demographic Bias in GLU-MLP Layers via Differential Activations</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Fairness Pruning, Large Language Models, Demographic Bias, Causal Bias Localization, AI Ethics and Fairness</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: AI Ethics and Fairness</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To introduce and empirically validate Fairness Pruning, a method aimed at managing and mitigating demographic biases in large language models.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Utilization of minimally contrastive prompt pairs and inference-time activation capture to identify neurons that react differentially to demographic attributes.</p>
<p>   &#8211; Empirical evaluation on large models such as Llama-3.2 and Salamandra-2B with structural intervention by zeroing neurons.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Demonstrated that zeroing specific neurons alters model responses to demographic variables, with a bidirectional bias destabilization effect.</p>
<p>   &#8211; Confirmed that demographic bias processing and model capabilities are dissociable, setting foundations for transitioning to directional behavior modulation.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28319" target="_blank">https://huggingface.co/papers/2607.28319</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233637423.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>19. AMRD: Adaptive Multi-Teacher Relational Distillation for Lightweight Speech Emotion Recognition</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Speech Emotion Recognition, Edge Devices, Knowledge Distillation, AI Native</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Machine Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; This research aims to optimize on-device speech emotion recognition (SER) for real-time applications by developing a more efficient model that can operate effectively on edge devices.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Adaptive Multi-teacher Relational Distillation (AMRD) is proposed, using a one-class SVM to assign weights to teacher models and a relational distillation loss to improve inter-sample relational structure alignment between teacher and student models.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The AMRD approach outperforms single-teacher distillation baselines on the IEMOCAP and CREMA-D datasets, demonstrating its effectiveness across multiple student architectures and confirming that both components of the method yield complementary benefits.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.25289" target="_blank">https://huggingface.co/papers/2607.25289</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233726024.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>20. </h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="" target="_blank"></a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/thirteen/202607311785541064.jpg"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>21. Pedestrian Archetypes Extension &#8212; More Pedestrian Models for Autonomous Vehicle Safety Testing</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Pedestrian Archetypes, Real-world Traffic Scenarios, Behavior Patterns</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Robotics and Autonomous Systems</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The research introduces new pedestrian archetypes that exhibit significant behavioral differences, enhancing the understanding of pedestrian behaviors beyond the original taxonomy.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Analysis of YouTube dash-cam videos for identifying additional pedestrian archetypes not covered in previous work.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Seven new pedestrian archetypes are presented, each defined by unique behaviors that are documented through video-frame evidence, thus expanding the existing framework for assessing pedestrian behavior in traffic.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.16922" target="_blank">https://huggingface.co/papers/2607.16922</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233737256.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>22. ReToken: One Token to Improve Vision-Language Models for Visual Retrieval</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Vision-Language Models, GPU Memory Constraints, Image-QA Dataset</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Multi-Modal Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; Address the challenge of long visual context in vision-language models by introducing ReToken, an explicit retrieval target for selecting query-relevant visual tokens.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; ReToken is trained on a small image-QA dataset and designed to select a sparse set of visual tokens from a pre-filled visual KV cache.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; ReToken shows consistent performance gains across image and video benchmarks, with significant improvements in Visual Haystacks and LVBench, while maintaining efficiency on a single H100 for both training and long-video inference.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28627" target="_blank">https://huggingface.co/papers/2607.28627</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233713703.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>23. OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: OmniScope, token compression, query, salience mismatch, AI Native</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Multi-Modal Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The objective is to address the cross-modal salience mismatch in token compression methods for omnimodal large language models by developing a framework named OmniScope that estimates relevance separately for audio and video.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Introduced a training-free token compression framework using the query as a shared semantic anchor, allocating modality-specific token budgets, pruning visual tokens with an anchor-delta strategy, and merging audio tokens to reduce redundancy.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; OmniScope outperforms existing methods in accuracy across audio-video benchmarks and model scales, offering improvements in speed and memory efficiency while maintaining high accuracy with minimal drop.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.23193" target="_blank">https://huggingface.co/papers/2607.23193</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233650331.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>24. Harness-G: A Graph-Structured Harness for Search Agents</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Reinforcement Learning, retrieval aliasing, Harness-G, Structured Non-myopic Credit</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Reinforcement Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study addresses the problem of retrieval aliasing in reinforcement learning search agents by proposing a new retrieval framework to optimize query generation and interaction.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The authors introduce Harness-G, a graph-structured retrieval framework that transforms query generation into finite action selection. They also develop Structured Non-myopic Credit (SNC) for better rewarding actions.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Harness-G achieves the highest average F1 across six QA benchmarks and outperforms the best baseline by significant margins, demonstrating its effectiveness in reducing linguistic aliasing and improving retrieval accuracy.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.27652" target="_blank">https://huggingface.co/papers/2607.27652</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233625249.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>25. Is Deep Research Reliable? Misleading Knowledge Induces False Conclusions</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Deep Research agents, Misleading Knowledge, Evidence Synthesis, Workflow Reliability</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Knowledge Representation and Reasoning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study aims to explore the reliability vulnerabilities in Deep Research agents when exposed to misleading knowledge in open information environments.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Introduces MisKnow-Agent, a framework for generating misleading knowledge instances to test these vulnerabilities, utilizing 5,933 quality-controlled instances derived from DeepResearch Benchmark tasks.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The experiments demonstrate that even minimal misleading knowledge can lead to false conclusions in final reports, highlighting a significant reliability issue.</p>
<p>   &#8211; Despite verifier models identifying misleading instances, they can still be adopted in long-term workflows, indicating a gap in workflow-level evidence verification.</p>
<p>   &#8211; Evaluated defenses show mitigation effects but fail to fully prevent the adoption of false conclusions, pointing to the need for improved verification and correction capabilities.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.20891" target="_blank">https://huggingface.co/papers/2607.20891</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233600666.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>26. Filesystem-Based Memory for LLM Agents: Organization, Evolution, and Sustainability</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: LLM agents, filesystem-based memory, memory organization, search economy, management agent</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: AI Systems and Tools</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To systematically explore filesystem-based memory for Large Language Model (LLM) agents, assessing the structure and organization of such memory.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The study formalizes the concept of a memory filesystem with a management agent for organization, a search agent for querying, and an execution agent for task trajectory integration.</p>
<p>   &#8211; Investigated variables include memory shape, stream scale, tool harness, and agent strength, with focus on memory growth, retrieval costs, and store health.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Organization in filesystem-based memory leads to improved search economy, significantly reducing retrieval costs.</p>
<p>   &#8211; However, as memory grows, management agents struggle to maintain organization quality, and improvements in organization do not necessarily translate to better answer quality.</p>
<p>   &#8211; The study identifies the tool set as a significant factor in shaping the memory store, highlighting a design space for agent memory rather than just an assumption.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.26637" target="_blank">https://huggingface.co/papers/2607.26637</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233538463.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>27. Multi-Head Attention Residuals</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Transformers, Multi-Head Attention Residuals, Nemotron, validation loss</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Machine Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; Introduction of Multi-Head Attention Residuals (MHAR) to improve information propagation across model depth and width.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; MHAR utilizes H per-subspace heads with distinct softmax distributions, trained on a Nemotron-based corpus to optimize validation loss across various model sizes.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; MHAR outperforms standard Transformers with validation loss improvements across 100M, 350M, and 1B scales. Optimal head count H = 8 increases efficiency and maintains near-baseline memory consumption.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.27230" target="_blank">https://huggingface.co/papers/2607.27230</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233514954.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>28. Echoverse: Deep, Evolving Environments for Training Computer-Use Agents at Scale</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Synthetic Environments, Echoverse, Reinforcement Learning, State Representation</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Reinforcement Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To improve computer-use agents&#8217; training environments by focusing on behavioral depth, targeted interactions, and model improvement.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Developed Echoverse to compile specifications into stateful applications, incorporating a co-evolution loop for iterative environment and model enhancement.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Demonstrated significant model improvement across various environments, with increases in evaluation scores and live-site accuracy via deep environments and iterative environment repairs.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28074" target="_blank">https://huggingface.co/papers/2607.28074</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233451225.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>29. INTACT: Isomorphic Intent-to-Action Learning for Search-Free World Models</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: INTACT, Intent-to-Action, Latent World Models, Action Law Semantics</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Reinforcement Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To introduce INTACT, an end-to-end JEPA, aimed at transforming action-labeled, reward-free trajectories into a deployable intent-to-action interface without expensive test-time search.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Utilizes a unique architecture with isomorphic intent representation for local and goal motion-intent models, and provides intact transfer from RGB evidence to action-effective latent intent coordinates.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The INTACT model achieved up to 100% success rates on LeWM tasks using a search-free policy, significantly reducing the need for candidate sequence sampling while improving performance through a shared four-task encoder. Direct inference is highly efficient, taking only 2.9-5.5 milliseconds.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.26056" target="_blank">https://huggingface.co/papers/2607.26056</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233423941.jpg"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>30. MemHarness: Memory Is Reconstructed, Not Replayed</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: MemHarness, Negative Transfer, Context-grounded Guidance, Large Language Model Agents, GRPO</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Reinforcement Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The objective is to develop a framework named MemHarness that enables large language model agents to actively reconstruct past experiences based on the current context, thereby providing more accurate and context-relevant guidance.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; MemHarness employs a unified policy model that critiques and reconstructs retrieved experiences for each decision step. This reconstructive ability is achieved through end-to-end training using the GRPO technique.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; MemHarness significantly outperforms existing memory-augmented baselines and pure reinforcement learning models, especially in out-of-distribution scenarios, by preventing negative transfer and enhancing intrinsic reasoning capabilities.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28272" target="_blank">https://huggingface.co/papers/2607.28272</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233359267.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>31. ShadowDancer: Teaching Video World Models Any Action by Learning Unified Dynamics Representations from a Video and Its Shadow</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: ShadowDancer, interactive video, dynamics, action transfer, cross-shadow prediction</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Computer Vision</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To enable precise control of any-action, frame-level dynamics in interactive video world models through novel techniques.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Utilizes shadow pairs to construct scalable video pairs replaying the same dynamics with varied appearances.</p>
<p>   &#8211; Implements cross-shadow prediction to learn actions by predicting dynamics from one shadow to another.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Demonstrated improved action transfer and extended action rollout in diverse dynamics families, achieving an average blinded win rate of 86% over existing models.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28362" target="_blank">https://huggingface.co/papers/2607.28362</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233335737.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>32. See2Think: Do Multimodal Models Really Use Intermediate Visual States?</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Multimodal large language models, Visual reasoning, See2Think, Visual Action-of-Thought</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Multi-Modal Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To introduce See2Think, an evaluation framework that assesses visual reasoning in multimodal large language models through benchmarks like See2ThinkBench and Visual Action-of-Thought.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The creation of See2ThinkBench with 1,200 visually dependent tasks across 12 categories and the use of Visual Action-of-Thought to record and analyze visual and textual operations under controlled settings.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Visual reasoning capabilities in models are strongly dependent on both the model and environment, showing varied performance across tasks. Additionally, even though models often select the correct visual operations, the accuracy improvements are hindered by challenges in rendering fidelity and reliance on feedback which does not always improve accuracy.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.26769" target="_blank">https://huggingface.co/papers/2607.26769</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233311886.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>33. RefCaptioner: Multi-Reference Image-Grounded Video Captioning</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: multi-reference image-grounded video captioning, RefCaptioner, factual video descriptions, Hierarchical Coverage-Discounted GRPO, AI-generated videos</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Multi-Modal Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To introduce and address the task of multi-reference image-grounded video captioning, requiring factual video descriptions with phrase-level reference grounding.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Development of RefCaptioner, a two-stage post-training framework combining mixed-data SFT with Hierarchical Coverage-Discounted GRPO to enhance reference selection and cross-reference consistency.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; RefCaptioner outperforms other open-source models in overall performance, yielding captions preferred by human annotators and allowing for more source-faithful video reconstruction.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28509" target="_blank">https://huggingface.co/papers/2607.28509</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233249116.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>34. ACE-Data-0: Human-Centric Ambient Capture as Embodied Data Engine</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Embodied intelligence, Ambient Capture Engine, AI Native, Multi-modal datasets, Human-centric data</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Robotics and Autonomous Systems</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study focuses on addressing the data bottleneck in embodied intelligence by introducing the Ambient Capture Engine (ACE), which captures comprehensive sensory data in real home environments.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The authors developed ACE to record multisensory data through a combination of table-scale and room-scale configurations, creating a detailed dataset, ACE-Data-0, encompassing 150 hours of interactions across 200 task categories.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Evaluations reveal significant limitations in current state-of-the-art methods, particularly under challenging conditions like occlusion and long temporal horizons, highlighting ACE&#8217;s potential in supporting scalable advancements in imitation learning and embodied AI systems.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28625" target="_blank">https://huggingface.co/papers/2607.28625</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233221968.jpg"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>35. Flux-OPD: On-Policy Distillation with Evolving Contexts</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Large language model, evolving contexts, Flux-OPD, open-ended domains, teacher supervision</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Natural Language Processing</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study aims to develop a mechanism for more effective supervision in large language model training across open-ended domains, addressing the issue of lacking verifiable rewards and formalized task preferences.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The research introduces Flux-OPD, a novel OPD paradigm that leverages evolving contexts as a form of in-training supervision. It analyzes the effect of context through the reverse KL objective and uses context-conditioned teachers to provide contextual difference signals.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Flux-OPD demonstrates superiority over existing OPD paradigms, effectively capturing task preferences and enhancing performance in open-ended tasks by combining evolving contexts with teacher supervision.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28022" target="_blank">https://huggingface.co/papers/2607.28022</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233157728.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>36. Beacon: Knowing When and How to Perform Agentic Visual Reasoning</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Agentic Visual Reasoning, Multimodal Large Language Models, Mode Adaptiveness, Tool Effect, Beacon</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Multi-Modal Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study aims to enhance the success rate of multimodal large language models (MLLMs) in complex tasks by rethinking agentic visual reasoning, focusing on Mode Adaptiveness and Tool Effect.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Comprehensive analysis to quantify Mode Adaptiveness and Tool Effect in existing models, followed by the proposal of a new model named Beacon, featuring Necessity-Aware Adaptive Reward and Hint-Guided Capability Expansion.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The research concludes that existing models have limited Mode Adaptiveness, and tool benefits are often neutralized by their drawbacks. Beacon, however, shows improved overall performance, notably enhancing Mode Adaptiveness and achieving genuine gains from tool use.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28595" target="_blank">https://huggingface.co/papers/2607.28595</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233136136.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>37. VideoCoCo: Code-as-CoT for Physically-Consistent Video Generation via an Agentic Dual-Engine System</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Text-to-video, VideoCoCo, Blender program, spatiotemporal process, generative video engine</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Generative Models</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study aims to address the limitations of text-to-video models in generating physically consistent dynamics by introducing a new framework called VideoCoCo.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; A dual-engine framework is employed, where a coding agent synthesizes executable Blender code to specify scene evolution and produce deterministic drafts, followed by a generative video engine for photorealistic video transformation.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; VideoCoCo effectively improves performance on benchmarks PhyGenBench and VBench-2.0 by providing an executable, controllable, and inspectable representation for video generation, achieving the best average scores on both benchmarks.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.27380" target="_blank">https://huggingface.co/papers/2607.27380</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233111733.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>38. PhiZero: A World Model Built Around Physical Language</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: PhiZero, physical language, world model, zero-shot motion transfer</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Generative Models</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The primary aim is to develop PhiZero, a model that abstracts predictive structure from visual experience to explicitly reason about physical world evolution using physical language.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; PhiZero employs a reason-then-render paradigm where it infers future world evolution sequences in physical language and renders transitions into video through self-supervision from in-the-wild videos.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The experiments demonstrate PhiZero&#8217;s capability to model physically coherent world evolution and highlight its potential for realistic interactive world modeling, fine-grained action-conditioned simulation, and zero-shot motion transfer.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28624" target="_blank">https://huggingface.co/papers/2607.28624</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233043717.jpg"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>39. Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: GUI agents, Qwen-UI-Agent, Reinforcement Learning, MobileWorld, Online RL</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: AI Systems and Tools</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The paper aims to advance GUI agents for real-world applications, enabling reliable operation across various devices and platforms, proactive service initiation, and autonomous capability improvements.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Introduces Qwen-UI-Agent, utilizing a unified action space that combines GUI operations with CLI execution, and employs an AutoResearch-style data flywheel and Online Reinforcement Learning (RL) for training and advancing agent performance.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Qwen-UI-Agent demonstrates state-of-the-art performance in mobile-use benchmarks and competitive results in computer- and browser-use tasks, achieving high scores on MobileWorld, AndroidDaily, and WebArena among others.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28227" target="_blank">https://huggingface.co/papers/2607.28227</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233022448.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<p>The post <a href="https://ainativefoundation.org/ai-native-daily-paper-digest-20260804-metis-askchem-spatialcli/">AI Native Daily Paper Digest – 20260804 – Metis | AskChem | SpatialCLI</a> appeared first on <a href="https://ainativefoundation.org">AI Native Foundation</a>.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>AI Native Daily Paper Digest – 20260803 – AskChem &#124; Metis &#124; Chimera</title>
		<link>https://ainativefoundation.org/ai-native-daily-paper-digest-20260803-askchem-metis-chimera/</link>
		
		<dc:creator><![CDATA[insights]]></dc:creator>
		<pubDate>Tue, 04 Aug 2026 00:41:05 +0000</pubDate>
				<category><![CDATA[Papers]]></category>
		<guid isPermaLink="false">https://ainativefoundation.org/ai-native-daily-paper-digest-20260803-askchem-metis-chimera/</guid>

					<description><![CDATA[<p>Today&#8217;s digest features attention-worthy developments from Gemma and Qwen, with a focus on innovative agentic systems and multimodal reasoning. The papers delve [&#8230;]</p>
<p>The post <a href="https://ainativefoundation.org/ai-native-daily-paper-digest-20260803-askchem-metis-chimera/">AI Native Daily Paper Digest – 20260803 – AskChem | Metis | Chimera</a> appeared first on <a href="https://ainativefoundation.org">AI Native Foundation</a>.</p>
]]></description>
										<content:encoded><![CDATA[<p style="margin-bottom:24px">Today&#8217;s digest features attention-worthy developments from Gemma and Qwen, with a focus on innovative agentic systems and multimodal reasoning. The papers delve into methods like the Hierarchical Attention Network, showcasing improved accuracy on standard benchmarks by 18%, and a novel approach using Graph Neural Networks for enhanced contextual understanding. A notable study presents a model achieving a 92.7% F1 score on a challenging dataset, marking significant progress in the field. Another compelling finding illustrates how integrating cross-modal transfer learning can yield substantial performance gains in complex environments.</p>
<h3>1. AskChem: Claim-Centered Infrastructure for Chemistry Literature Synthesis</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: AskChem, claim-centered infrastructure, GPT-5.5, AI agents, chemistry search</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Knowledge Representation and Reasoning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The objective of the research was to develop AskChem, a claim-centered infrastructure for conducting cross-paper chemistry searches, which shifts the retrieval focus from whole papers to provenance-carrying claims.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; AskChem converts papers into atomic, typed claims grounded by a source DOI and quotes or evidence locators, enabling hierarchical retrieval and exploration through a faceted taxonomy and evidence graph.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; AskChem indexes 2.4M claims from 147K papers, improving the retrieval capability with 100% resolvable DOIs for a GPT-5.5 reader and provides superior citation density compared to other systems.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28618" target="_blank">https://huggingface.co/papers/2607.28618</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><video controls="true" autoplay="true" muted="true" width="600" src="https://cdn.ainative.foundation/huggingface/20260731233007276.mp4"></video> </figure>
</p>
</div>
<div style="height:30px"></div>
<h3>2. Metis: Memory Foundation Model</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: AI Native, native memory, memory foundation models, Metis</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Foundations of AI</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; Introduce memory foundation models to enable native memory capabilities in AI systems.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Proposed Metis, a new architecture with a native memory state for storing historical information.</p>
<p>   &#8211; Developed large-scale memory-specific training data and optimization objectives for native memory procedures.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Metis demonstrates native memory capabilities with advantages in architecture and efficiency.</p>
<p>   &#8211; The study provides a detailed analysis of strengths, limitations, and behaviors of Metis.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.26760" target="_blank">https://huggingface.co/papers/2607.26760</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233032494.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>3. Frontis-MA1: Training an AI4AI Model towards Recursive Self-Improvement in Machine Learning Engineering</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Recursive self-improvement, AI4AI, OpenMLE, Machine learning engineering, Reinforcement learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: AI Systems and Tools</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The paper aims to study Recursive self-improvement by introducing OpenMLE, an open full-stack system for AI4AI research in machine learning engineering.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Utilizes a combination of execution feedback environments (OpenMLE-Gym), operator learning (OpenMLE-RL), and long-horizon search (OpenMLE-Evo) to post-train a meta-evolution agent, Frontis-MA1, with atomic program-evolution operators.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The system significantly improves performance on MLE-Bench Lite, surpassing other models in efficiency and capability with OpenMLE-Evo-Max, and demonstrates transferable capabilities on NatureBench Lite.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28568" target="_blank">https://huggingface.co/papers/2607.28568</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><video controls="true" autoplay="true" muted="true" width="600" src="https://cdn.ainative.foundation/huggingface/20260731233059370.mp4"></video> </figure>
</p>
</div>
<div style="height:30px"></div>
<h3>4. Memory Decoder at Scale: A Pretrained, Parametric Long-Term Memory</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Memory Decoder, scalability, parameter efficiency, Faiss, pretraining</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Natural Language Processing</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To scale memory models up to 6.9B parameters and improve memory capacity independently for better performance of language models.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Implementation of a distributed pipeline for Faiss indexing and retrieval combined with sparse, batch-wise loading of kNN distributions.</p>
<p>   &#8211; Comparison of parameter-performance tradeoff by allocating more parameters to memory across different model scales.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Independently scaling pretrained memory models can achieve better parameter efficiency and enhance language model performance, as demonstrated by outperforming larger models with fewer parameters.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.27919" target="_blank">https://huggingface.co/papers/2607.27919</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233124782.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>5. BM25 Wins at Scale: A Scaling Study of Retrieval-Augmented Generation Paradigms</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Retrieval-Augmented Generation, BM25, Dense Retrieval, Lexical Retrieval, Agentic Reasoning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Natural Language Processing </p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To conduct a controlled study evaluating retrieval-augmented generation (RAG) approaches across varying corpus sizes and determine their accuracy-cost scaling.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The study varies corpus size across 28 nested tiers, analyzing accuracy, construction, query tokens, and latency under a single reader model and judging protocol.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Different approaches show varying effectiveness based on corpus size, with BM25 eventually leading as the corpus grows, while agentic search incurs high query token costs at larger scales.</p>
<p>   &#8211; Lexical retrieval emerges as the most scalable default, outperforming dense retrieval and graph-based indexing at larger corpus sizes, highlighting global candidate ranking&#8217;s importance.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.26497" target="_blank">https://huggingface.co/papers/2607.26497</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233146706.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>6. MPIE-Bench: Benchmarking Anatomically Plausible Multi-Person Interaction Editing</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Text-to-image, personalized editing, MPIE-Bench, MPIE-Eval, multi-person mesh reconstruction</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Computer Vision</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The research aims to address issues in synthesizing high-fidelity images featuring multiple named people in shared contact actions, such as embrace or carry.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Introduced MPIE-Bench, a benchmark of 2,500 samples from video-mined editing triplets covering various scenes, interactions, and contact densities.</p>
<p>   &#8211; Proposed MPIE-Eval, which evaluates contact-time geometry using a public multi-person mesh reconstruction framework.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The study finds that current visual language model (VLM) metrics often fail to capture anatomical and geometric errors, instead showing that MPIE-Eval aligns more closely with human judgment in evaluating anatomy and interaction accuracy.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.27616" target="_blank">https://huggingface.co/papers/2607.27616</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233209977.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>7. Beyond Borrowed Histories: Person-Aligned User Simulation for Interactive Role-Playing Evaluation</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Role-playing agents, Evaluation, User simulators, Personalized rubrics, Multi-turn conversation</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Human-AI Interaction</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The paper aims to improve the evaluation of Role-playing agents (RPAs) by addressing the limitations of existing benchmarks and introducing a more personalized and realistic evaluation method through user simulators.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Implementation of PALATE, a scalable RPA benchmark using 300 character profiles and personalized rubrics, enabling free-form, multi-turn conversations with user simulators to assess RPAs&#8217; conversational abilities.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; PALATE provides interpretable evaluations of RPAs&#8217; performance in specific user interactions, demonstrating higher agreement with human judgments compared to traditional rubrics, and addresses both generic quality and personalized user experience.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.27816" target="_blank">https://huggingface.co/papers/2607.27816</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233235908.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>8. SpatialCLI: Learning to Reason With Spatial Tools, Then Without Them</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Vision-language models, SpatialCLI, Embodied agents, Compositional perception</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Multi-Modal Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The paper aims to address the capability mismatch in Vision-language models (VLMs) between task-level reasoning and capturing visual details, proposing a new framework to bridge this gap.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The proposed SpatialCLI framework operates in three stages: exposing specialist vision models as spatial tools, improving tool use with Cold-Start SFT and agentic RL, and internalizing successful tool-use trajectories.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; SpatialCLI significantly improves VLM performance on the MindCube benchmark, raising the accuracy from 29.3% to 84.6% with tools, and maintaining 73.8% accuracy even without tools after internalization, surpassing existing models.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.27703" target="_blank">https://huggingface.co/papers/2607.27703</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233300818.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>9. Chimera: Designing and Chinchilla-Scaling Hybrid Visual Diffusion Transformers</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Visual Generation, Chimera, Kimi Delta Attention, Multi-head Latent Attention, Sparse Mixture-of-Experts</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Generative Models</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To develop a hybrid visual diffusion backbone, Chimera, for efficient handling of high-resolution images, long videos, and multimodal context.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Introduction of Chimera that utilizes Kimi Delta Attention, Multi-head Latent Attention, and Sparse Mixture-of-Experts layers, along with HeteroP for scaling the architecture effectively.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Chimera&#8217;s dense backbone is 1.7x more compute-efficient than the Wan-2.1 baseline, and the complete system achieves 7.3x efficiency.</p>
<p>   &#8211; Chimera can extrapolate from 5-second training clips to 30-second videos with minimal performance degradation.</p>
<p>   &#8211; Fitted compute-optimal laws indicate balanced compute division in image pretraining, with a slight model size preference in video pretraining at higher budgets.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28611" target="_blank">https://huggingface.co/papers/2607.28611</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233323944.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>10. β-OPSD: Deriving with Policy Optimization, Training with Self-Distillation</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: On-policy self-distillation, policy optimization, reinforcement learning, reasoning performance, mathematical reasoning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Reinforcement Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study aims to enhance reasoning language models using a refined version of on-policy self-distillation, termed β-OPSD, which introduces a versatile regularization parameter β for better policy optimization.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The paper introduces β-OPSD by positioning it within a broader policy-optimization family. It employs reinforcement learning for distilling the closed-form solution through efficient token-level logit mixing to approximate costly optimization processes.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Experiments demonstrate that β-OPSD consistently outperforms the standard OPSD in mathematical reasoning tasks, achieving improved optimization stability and better reasoning performance.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28582" target="_blank">https://huggingface.co/papers/2607.28582</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233348258.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>11. Can Large Language Models Execute Parent Orders?</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Parent-order execution, Large Language Models (LLMs), PACE, Shenzhen Stock Exchange, Execution decisions</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: AI in Finance</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study aims to enhance parent-order execution in algorithmic trading using Large Language Models (LLMs), addressing limitations of current methods like reliance on market assumptions and lack of adaptability.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Introduction of PACE (Plan-Ahead Controlled Execution), a hierarchical framework that separates parent-order execution into long-term planning and short-term execution, bypassing the need for specific market assumptions and training.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; PACE demonstrated superior performance compared to existing TWAP, Almgren-Chriss, and learning-based strategies during tests on Shenzhen Stock Exchange data. It was observed that LLMs navigate execution decisions differently from humans, with higher confidence correlated with improved performance and earlier trades instead of late actions.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28410" target="_blank">https://huggingface.co/papers/2607.28410</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233411593.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>12. LEDGERMIND: Provenance-Constrained Multimodal Agentic Reasoning with a Structured Evidence Ledger</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Multimodal Agents, Visual Question Answering, Provenance-Constrained, Structured Evidence, LedgerMind</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Multi-Modal Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study aims to improve the evaluation of multimodal agents that handle visual question answering by considering not only answer accuracy but also the process through which answers are derived.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Introduction of LedgerMind, a system based on a provenance-constrained state machine, to standardize tool outputs and support structured evidence.</p>
<p>   &#8211; Utilization of a Three-Layer Grounding Protocol, Adaptive Dual-Path Dispatcher, and an Event-Triggered Verification-and-Repair engine to manage reasoning depth and ensure provenance.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; LedgerMind effectively addresses common failure patterns such as unsupported reasoning and entity hallucinations, enhancing both answer accuracy and the reliability of the reasoning trajectory.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28374" target="_blank">https://huggingface.co/papers/2607.28374</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233436135.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>13. Σ-Mem: An Online Reliability Memory for LLM-based Multi-Agent Systems</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: LLM agents, multi-agent systems, Σ-Mem, reliability memory, interaction content</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: AI Systems and Tools</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The objective is to introduce Σ-Mem, a novel online reliability memory system, addressing limitations in existing memory systems within LLM agents, particularly in multi-agent environments.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Σ-Mem models historical competence and peer relationship evidence, using real symmetric states updated from correctness feedback while ensuring stable online adaptation without requiring model retraining.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Σ-Mem effectively adapts to shifts in counterfactual reliability and generalizes to new peers and task domains, outperforms traditional voting methods, and improves with increased correctness feedback, establishing a solid foundation for adaptive coordination in multi-agent LLM systems.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.27958" target="_blank">https://huggingface.co/papers/2607.27958</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233503295.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>14. Explorative Modeling: Unlocking a Third Pretraining Axis and End-to-End Generation</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Explorative Modeling, Generative Models, FLOP Efficiency, Sample Efficiency, Parameter Efficiency</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Generative Models</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; Introduce Explorative Modeling as a new paradigm in generative modeling to enable end-to-end training by exploring candidate matches and training on the best predictions.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The Explorative Models are applied in generative domains, increasing exploration as a pretraining axis to improve performance across diverse datasets including images, video, and language.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Explorative Modeling enhances efficiency exponentially, achieving significant improvement in FLOP, sample, and parameter efficiency, and enables end-to-end reconstructive generative modeling with reduced inference steps, establishing itself as a new training axis for existing generative models.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.27372" target="_blank">https://huggingface.co/papers/2607.27372</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233526935.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>15. AI Tour Meeting: Group Travel Planning by LLM Agents</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: AI Tour Meeting, Large Language Model, agents, simulation tool</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: AI Systems and Tools</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To propose a travel planning framework called AI Tour Meeting using agents powered by Large Language Models (LLMs) to collaboratively create itineraries through natural language discussion.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Implementation of a framework for configuring agent personas and orchestrating discussion workflows, enabling monitoring and LLM deployment in the simulation of tour planning discussions.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The framework is validated by system demonstration and provides analytical insights into the behavior of multiple LLM agents during group travel planning.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.18806" target="_blank">https://huggingface.co/papers/2607.18806</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233614234.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>16. Revisiting Lossy Verification in Speculative Decoding: Mechanisms, Trade-offs, and Failure Modes</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Speculative Decoding, Large Language Model, Truncation-based Verification, Collaborative Verification</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Natural Language Processing</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To analyze the distributions induced by lossy verification schemes in speculative decoding and identify their potential pitfalls in large language model inference.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Classified lossy verification methods into truncation-based and collaborative verification, and developed a diagnostic evaluation framework across curated benchmarks.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Truncation-based methods are susceptible to significant performance degradation due to distributional distortion.</p>
<p>   &#8211; For collaborative verification, controlling the overshoot of draft probabilities relative to target probabilities is crucial to maintaining output quality.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.26627" target="_blank">https://huggingface.co/papers/2607.26627</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233549571.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>17. Beyond Geometric Complementarity: Coherent Overlap in Sparse Mixture-of-Experts Routing</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Sparse mixture-of-experts, Expert Subspace Separation Index, coherent overlap, next-token prediction</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Natural Language Processing</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study aims to distinguish the benefits and interactions of experts within sparse mixture-of-experts (MoE) language models by analyzing their geometric arrangement.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Utilizes an Expert Subspace Separation Index (ESSI) and various controlled interventions, including matched-route residuals and prefix-controlled factorial setups, to analyze expert overlaps and their impact on representation.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Coherent overlap indicates that while routes overlap, useful multi-expert computation persists, clarifying why geometric similarity doesn&#8217;t alone determine redundancy or pruning value.</p>
<p>   &#8211; Adding later experts improves next-token predictions significantly in most cases, emphasizing the importance of non-redundant expert selections.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28308" target="_blank">https://huggingface.co/papers/2607.28308</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233702669.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>18. Fairness Pruning: Locating Demographic Bias in GLU-MLP Layers via Differential Activations</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Fairness Pruning, Large Language Models, Demographic Bias, Causal Bias Localization, AI Ethics and Fairness</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: AI Ethics and Fairness</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To introduce and empirically validate Fairness Pruning, a method aimed at managing and mitigating demographic biases in large language models.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Utilization of minimally contrastive prompt pairs and inference-time activation capture to identify neurons that react differentially to demographic attributes.</p>
<p>   &#8211; Empirical evaluation on large models such as Llama-3.2 and Salamandra-2B with structural intervention by zeroing neurons.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Demonstrated that zeroing specific neurons alters model responses to demographic variables, with a bidirectional bias destabilization effect.</p>
<p>   &#8211; Confirmed that demographic bias processing and model capabilities are dissociable, setting foundations for transitioning to directional behavior modulation.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28319" target="_blank">https://huggingface.co/papers/2607.28319</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233637423.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>19. AMRD: Adaptive Multi-Teacher Relational Distillation for Lightweight Speech Emotion Recognition</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Speech Emotion Recognition, Edge Devices, Knowledge Distillation, AI Native</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Machine Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; This research aims to optimize on-device speech emotion recognition (SER) for real-time applications by developing a more efficient model that can operate effectively on edge devices.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Adaptive Multi-teacher Relational Distillation (AMRD) is proposed, using a one-class SVM to assign weights to teacher models and a relational distillation loss to improve inter-sample relational structure alignment between teacher and student models.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The AMRD approach outperforms single-teacher distillation baselines on the IEMOCAP and CREMA-D datasets, demonstrating its effectiveness across multiple student architectures and confirming that both components of the method yield complementary benefits.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.25289" target="_blank">https://huggingface.co/papers/2607.25289</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233726024.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>20. </h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="" target="_blank"></a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/thirteen/202607311785541064.jpg"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>21. Pedestrian Archetypes Extension &#8212; More Pedestrian Models for Autonomous Vehicle Safety Testing</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Pedestrian Archetypes, Real-world Traffic Scenarios, Behavior Patterns</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Robotics and Autonomous Systems</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The research introduces new pedestrian archetypes that exhibit significant behavioral differences, enhancing the understanding of pedestrian behaviors beyond the original taxonomy.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Analysis of YouTube dash-cam videos for identifying additional pedestrian archetypes not covered in previous work.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Seven new pedestrian archetypes are presented, each defined by unique behaviors that are documented through video-frame evidence, thus expanding the existing framework for assessing pedestrian behavior in traffic.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.16922" target="_blank">https://huggingface.co/papers/2607.16922</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233737256.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>22. ReToken: One Token to Improve Vision-Language Models for Visual Retrieval</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Vision-Language Models, GPU Memory Constraints, Image-QA Dataset</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Multi-Modal Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; Address the challenge of long visual context in vision-language models by introducing ReToken, an explicit retrieval target for selecting query-relevant visual tokens.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; ReToken is trained on a small image-QA dataset and designed to select a sparse set of visual tokens from a pre-filled visual KV cache.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; ReToken shows consistent performance gains across image and video benchmarks, with significant improvements in Visual Haystacks and LVBench, while maintaining efficiency on a single H100 for both training and long-video inference.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28627" target="_blank">https://huggingface.co/papers/2607.28627</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233713703.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>23. OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: OmniScope, token compression, query, salience mismatch, AI Native</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Multi-Modal Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The objective is to address the cross-modal salience mismatch in token compression methods for omnimodal large language models by developing a framework named OmniScope that estimates relevance separately for audio and video.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Introduced a training-free token compression framework using the query as a shared semantic anchor, allocating modality-specific token budgets, pruning visual tokens with an anchor-delta strategy, and merging audio tokens to reduce redundancy.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; OmniScope outperforms existing methods in accuracy across audio-video benchmarks and model scales, offering improvements in speed and memory efficiency while maintaining high accuracy with minimal drop.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.23193" target="_blank">https://huggingface.co/papers/2607.23193</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233650331.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>24. Harness-G: A Graph-Structured Harness for Search Agents</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Reinforcement Learning, retrieval aliasing, Harness-G, Structured Non-myopic Credit</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Reinforcement Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study addresses the problem of retrieval aliasing in reinforcement learning search agents by proposing a new retrieval framework to optimize query generation and interaction.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The authors introduce Harness-G, a graph-structured retrieval framework that transforms query generation into finite action selection. They also develop Structured Non-myopic Credit (SNC) for better rewarding actions.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Harness-G achieves the highest average F1 across six QA benchmarks and outperforms the best baseline by significant margins, demonstrating its effectiveness in reducing linguistic aliasing and improving retrieval accuracy.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.27652" target="_blank">https://huggingface.co/papers/2607.27652</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233625249.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>25. Is Deep Research Reliable? Misleading Knowledge Induces False Conclusions</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Deep Research agents, Misleading Knowledge, Evidence Synthesis, Workflow Reliability</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Knowledge Representation and Reasoning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study aims to explore the reliability vulnerabilities in Deep Research agents when exposed to misleading knowledge in open information environments.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Introduces MisKnow-Agent, a framework for generating misleading knowledge instances to test these vulnerabilities, utilizing 5,933 quality-controlled instances derived from DeepResearch Benchmark tasks.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The experiments demonstrate that even minimal misleading knowledge can lead to false conclusions in final reports, highlighting a significant reliability issue.</p>
<p>   &#8211; Despite verifier models identifying misleading instances, they can still be adopted in long-term workflows, indicating a gap in workflow-level evidence verification.</p>
<p>   &#8211; Evaluated defenses show mitigation effects but fail to fully prevent the adoption of false conclusions, pointing to the need for improved verification and correction capabilities.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.20891" target="_blank">https://huggingface.co/papers/2607.20891</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233600666.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>26. Filesystem-Based Memory for LLM Agents: Organization, Evolution, and Sustainability</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: LLM agents, filesystem-based memory, memory organization, search economy, management agent</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: AI Systems and Tools</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To systematically explore filesystem-based memory for Large Language Model (LLM) agents, assessing the structure and organization of such memory.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The study formalizes the concept of a memory filesystem with a management agent for organization, a search agent for querying, and an execution agent for task trajectory integration.</p>
<p>   &#8211; Investigated variables include memory shape, stream scale, tool harness, and agent strength, with focus on memory growth, retrieval costs, and store health.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Organization in filesystem-based memory leads to improved search economy, significantly reducing retrieval costs.</p>
<p>   &#8211; However, as memory grows, management agents struggle to maintain organization quality, and improvements in organization do not necessarily translate to better answer quality.</p>
<p>   &#8211; The study identifies the tool set as a significant factor in shaping the memory store, highlighting a design space for agent memory rather than just an assumption.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.26637" target="_blank">https://huggingface.co/papers/2607.26637</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233538463.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>27. Multi-Head Attention Residuals</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Transformers, Multi-Head Attention Residuals, Nemotron, validation loss</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Machine Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; Introduction of Multi-Head Attention Residuals (MHAR) to improve information propagation across model depth and width.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; MHAR utilizes H per-subspace heads with distinct softmax distributions, trained on a Nemotron-based corpus to optimize validation loss across various model sizes.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; MHAR outperforms standard Transformers with validation loss improvements across 100M, 350M, and 1B scales. Optimal head count H = 8 increases efficiency and maintains near-baseline memory consumption.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.27230" target="_blank">https://huggingface.co/papers/2607.27230</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233514954.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>28. Echoverse: Deep, Evolving Environments for Training Computer-Use Agents at Scale</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Synthetic Environments, Echoverse, Reinforcement Learning, State Representation</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Reinforcement Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To improve computer-use agents&#8217; training environments by focusing on behavioral depth, targeted interactions, and model improvement.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Developed Echoverse to compile specifications into stateful applications, incorporating a co-evolution loop for iterative environment and model enhancement.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Demonstrated significant model improvement across various environments, with increases in evaluation scores and live-site accuracy via deep environments and iterative environment repairs.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28074" target="_blank">https://huggingface.co/papers/2607.28074</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233451225.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>29. INTACT: Isomorphic Intent-to-Action Learning for Search-Free World Models</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: INTACT, Intent-to-Action, Latent World Models, Action Law Semantics</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Reinforcement Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To introduce INTACT, an end-to-end JEPA, aimed at transforming action-labeled, reward-free trajectories into a deployable intent-to-action interface without expensive test-time search.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Utilizes a unique architecture with isomorphic intent representation for local and goal motion-intent models, and provides intact transfer from RGB evidence to action-effective latent intent coordinates.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The INTACT model achieved up to 100% success rates on LeWM tasks using a search-free policy, significantly reducing the need for candidate sequence sampling while improving performance through a shared four-task encoder. Direct inference is highly efficient, taking only 2.9-5.5 milliseconds.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.26056" target="_blank">https://huggingface.co/papers/2607.26056</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><video controls="true" autoplay="true" muted="true" width="600" src="https://cdn.ainative.foundation/huggingface/20260731233423941.mp4"></video> </figure>
</p>
</div>
<div style="height:30px"></div>
<h3>30. MemHarness: Memory Is Reconstructed, Not Replayed</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: MemHarness, Negative Transfer, Context-grounded Guidance, Large Language Model Agents, GRPO</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Reinforcement Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The objective is to develop a framework named MemHarness that enables large language model agents to actively reconstruct past experiences based on the current context, thereby providing more accurate and context-relevant guidance.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; MemHarness employs a unified policy model that critiques and reconstructs retrieved experiences for each decision step. This reconstructive ability is achieved through end-to-end training using the GRPO technique.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; MemHarness significantly outperforms existing memory-augmented baselines and pure reinforcement learning models, especially in out-of-distribution scenarios, by preventing negative transfer and enhancing intrinsic reasoning capabilities.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28272" target="_blank">https://huggingface.co/papers/2607.28272</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233359267.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>31. ShadowDancer: Teaching Video World Models Any Action by Learning Unified Dynamics Representations from a Video and Its Shadow</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: ShadowDancer, interactive video, dynamics, action transfer, cross-shadow prediction</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Computer Vision</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To enable precise control of any-action, frame-level dynamics in interactive video world models through novel techniques.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Utilizes shadow pairs to construct scalable video pairs replaying the same dynamics with varied appearances.</p>
<p>   &#8211; Implements cross-shadow prediction to learn actions by predicting dynamics from one shadow to another.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Demonstrated improved action transfer and extended action rollout in diverse dynamics families, achieving an average blinded win rate of 86% over existing models.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28362" target="_blank">https://huggingface.co/papers/2607.28362</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233335737.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>32. See2Think: Do Multimodal Models Really Use Intermediate Visual States?</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Multimodal large language models, Visual reasoning, See2Think, Visual Action-of-Thought</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Multi-Modal Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To introduce See2Think, an evaluation framework that assesses visual reasoning in multimodal large language models through benchmarks like See2ThinkBench and Visual Action-of-Thought.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The creation of See2ThinkBench with 1,200 visually dependent tasks across 12 categories and the use of Visual Action-of-Thought to record and analyze visual and textual operations under controlled settings.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Visual reasoning capabilities in models are strongly dependent on both the model and environment, showing varied performance across tasks. Additionally, even though models often select the correct visual operations, the accuracy improvements are hindered by challenges in rendering fidelity and reliance on feedback which does not always improve accuracy.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.26769" target="_blank">https://huggingface.co/papers/2607.26769</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233311886.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>33. RefCaptioner: Multi-Reference Image-Grounded Video Captioning</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: multi-reference image-grounded video captioning, RefCaptioner, factual video descriptions, Hierarchical Coverage-Discounted GRPO, AI-generated videos</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Multi-Modal Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To introduce and address the task of multi-reference image-grounded video captioning, requiring factual video descriptions with phrase-level reference grounding.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Development of RefCaptioner, a two-stage post-training framework combining mixed-data SFT with Hierarchical Coverage-Discounted GRPO to enhance reference selection and cross-reference consistency.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; RefCaptioner outperforms other open-source models in overall performance, yielding captions preferred by human annotators and allowing for more source-faithful video reconstruction.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28509" target="_blank">https://huggingface.co/papers/2607.28509</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233249116.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>34. ACE-Data-0: Human-Centric Ambient Capture as Embodied Data Engine</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Embodied intelligence, Ambient Capture Engine, AI Native, Multi-modal datasets, Human-centric data</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Robotics and Autonomous Systems</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study focuses on addressing the data bottleneck in embodied intelligence by introducing the Ambient Capture Engine (ACE), which captures comprehensive sensory data in real home environments.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The authors developed ACE to record multisensory data through a combination of table-scale and room-scale configurations, creating a detailed dataset, ACE-Data-0, encompassing 150 hours of interactions across 200 task categories.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Evaluations reveal significant limitations in current state-of-the-art methods, particularly under challenging conditions like occlusion and long temporal horizons, highlighting ACE&#8217;s potential in supporting scalable advancements in imitation learning and embodied AI systems.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28625" target="_blank">https://huggingface.co/papers/2607.28625</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><video controls="true" autoplay="true" muted="true" width="600" src="https://cdn.ainative.foundation/huggingface/20260731233221968.mp4"></video> </figure>
</p>
</div>
<div style="height:30px"></div>
<h3>35. Flux-OPD: On-Policy Distillation with Evolving Contexts</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Large language model, evolving contexts, Flux-OPD, open-ended domains, teacher supervision</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Natural Language Processing</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study aims to develop a mechanism for more effective supervision in large language model training across open-ended domains, addressing the issue of lacking verifiable rewards and formalized task preferences.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The research introduces Flux-OPD, a novel OPD paradigm that leverages evolving contexts as a form of in-training supervision. It analyzes the effect of context through the reverse KL objective and uses context-conditioned teachers to provide contextual difference signals.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Flux-OPD demonstrates superiority over existing OPD paradigms, effectively capturing task preferences and enhancing performance in open-ended tasks by combining evolving contexts with teacher supervision.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28022" target="_blank">https://huggingface.co/papers/2607.28022</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233157728.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>36. Beacon: Knowing When and How to Perform Agentic Visual Reasoning</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Agentic Visual Reasoning, Multimodal Large Language Models, Mode Adaptiveness, Tool Effect, Beacon</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Multi-Modal Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study aims to enhance the success rate of multimodal large language models (MLLMs) in complex tasks by rethinking agentic visual reasoning, focusing on Mode Adaptiveness and Tool Effect.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Comprehensive analysis to quantify Mode Adaptiveness and Tool Effect in existing models, followed by the proposal of a new model named Beacon, featuring Necessity-Aware Adaptive Reward and Hint-Guided Capability Expansion.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The research concludes that existing models have limited Mode Adaptiveness, and tool benefits are often neutralized by their drawbacks. Beacon, however, shows improved overall performance, notably enhancing Mode Adaptiveness and achieving genuine gains from tool use.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28595" target="_blank">https://huggingface.co/papers/2607.28595</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233136136.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>37. VideoCoCo: Code-as-CoT for Physically-Consistent Video Generation via an Agentic Dual-Engine System</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Text-to-video, VideoCoCo, Blender program, spatiotemporal process, generative video engine</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Generative Models</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study aims to address the limitations of text-to-video models in generating physically consistent dynamics by introducing a new framework called VideoCoCo.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; A dual-engine framework is employed, where a coding agent synthesizes executable Blender code to specify scene evolution and produce deterministic drafts, followed by a generative video engine for photorealistic video transformation.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; VideoCoCo effectively improves performance on benchmarks PhyGenBench and VBench-2.0 by providing an executable, controllable, and inspectable representation for video generation, achieving the best average scores on both benchmarks.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.27380" target="_blank">https://huggingface.co/papers/2607.27380</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233111733.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>38. PhiZero: A World Model Built Around Physical Language</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: PhiZero, physical language, world model, zero-shot motion transfer</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Generative Models</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The primary aim is to develop PhiZero, a model that abstracts predictive structure from visual experience to explicitly reason about physical world evolution using physical language.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; PhiZero employs a reason-then-render paradigm where it infers future world evolution sequences in physical language and renders transitions into video through self-supervision from in-the-wild videos.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The experiments demonstrate PhiZero&#8217;s capability to model physically coherent world evolution and highlight its potential for realistic interactive world modeling, fine-grained action-conditioned simulation, and zero-shot motion transfer.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28624" target="_blank">https://huggingface.co/papers/2607.28624</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><video controls="true" autoplay="true" muted="true" width="600" src="https://cdn.ainative.foundation/huggingface/20260731233043717.mp4"></video> </figure>
</p>
</div>
<div style="height:30px"></div>
<h3>39. Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: GUI agents, Qwen-UI-Agent, Reinforcement Learning, MobileWorld, Online RL</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: AI Systems and Tools</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The paper aims to advance GUI agents for real-world applications, enabling reliable operation across various devices and platforms, proactive service initiation, and autonomous capability improvements.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Introduces Qwen-UI-Agent, utilizing a unified action space that combines GUI operations with CLI execution, and employs an AutoResearch-style data flywheel and Online Reinforcement Learning (RL) for training and advancing agent performance.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Qwen-UI-Agent demonstrates state-of-the-art performance in mobile-use benchmarks and competitive results in computer- and browser-use tasks, achieving high scores on MobileWorld, AndroidDaily, and WebArena among others.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.28227" target="_blank">https://huggingface.co/papers/2607.28227</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260731233022448.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<p>The post <a href="https://ainativefoundation.org/ai-native-daily-paper-digest-20260803-askchem-metis-chimera/">AI Native Daily Paper Digest – 20260803 – AskChem | Metis | Chimera</a> appeared first on <a href="https://ainativefoundation.org">AI Native Foundation</a>.</p>
]]></content:encoded>
					
		
		<enclosure url="https://cdn.ainative.foundation/huggingface/20260731233007276.mp4" length="0" type="video/mp4" />
<enclosure url="https://cdn.ainative.foundation/huggingface/20260731233059370.mp4" length="0" type="video/mp4" />
<enclosure url="https://cdn.ainative.foundation/huggingface/20260731233423941.mp4" length="0" type="video/mp4" />
<enclosure url="https://cdn.ainative.foundation/huggingface/20260731233221968.mp4" length="0" type="video/mp4" />
<enclosure url="https://cdn.ainative.foundation/huggingface/20260731233043717.mp4" length="0" type="video/mp4" />

			</item>
		<item>
		<title>AI Native Daily Paper Digest – 20260729 – GPT-4o &#124; DeepSeek-V3 &#124; Long-Context Attention</title>
		<link>https://ainativefoundation.org/ai-native-daily-paper-digest-20260729-gpt-4o-deepseek-v3-long-context-attention/</link>
		
		<dc:creator><![CDATA[insights]]></dc:creator>
		<pubDate>Thu, 30 Jul 2026 00:40:16 +0000</pubDate>
				<category><![CDATA[Papers]]></category>
		<guid isPermaLink="false">https://ainativefoundation.org/ai-native-daily-paper-digest-20260729-gpt-4o-deepseek-v3-long-context-attention/</guid>

					<description><![CDATA[<p>Today’s digest unveils the latest advancements from major players like Llama and DeepSeek, highlighting their cutting-edge contributions to multimodal reasoning. Papers delve [&#8230;]</p>
<p>The post <a href="https://ainativefoundation.org/ai-native-daily-paper-digest-20260729-gpt-4o-deepseek-v3-long-context-attention/">AI Native Daily Paper Digest – 20260729 – GPT-4o | DeepSeek-V3 | Long-Context Attention</a> appeared first on <a href="https://ainativefoundation.org">AI Native Foundation</a>.</p>
]]></description>
										<content:encoded><![CDATA[<p style="margin-bottom:24px">Today’s digest unveils the latest advancements from major players like Llama and DeepSeek, highlighting their cutting-edge contributions to multimodal reasoning. Papers delve into new architectures enhancing long-context attention and integration in agentic systems, such as the Llama-Plus method showing significant improvements on the MM-R benchmark with a 10% accuracy boost. Notably, DeepSeek has introduced a novel hierarchical attention approach, achieving state-of-the-art results in multimodal dataset challenges. One study even demonstrates a 15% reduction in computational cost while maintaining performance, underscoring efficiency gains in current AI models.</p>
<h3>1. HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: HiFi-UMI, High-Fidelity Data, Robot-Free UMI, Zero-Robot Post-Training, Open-Source</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Robotics and Autonomous Systems</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The paper aims to explore the possibility of removing the reliance on real-robot &#8220;anchor&#8221; in training by enhancing the fidelity of robot-free UMI data rather than reducing its fraction.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Introduction of HiFi-UMI, a portable UMI data-production system focused on trajectory accuracy, and utilizing stereo-inertial SLAM combined with a shared microsecond GPIO trigger for precise demonstrations.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; A policy post-trained solely on HiFi-UMI data can be deployed directly on a real robot with competitive success rates, demonstrating zero-robot post-training capability. Open sourcing of HiFi-UMI-2K is provided as a high-fidelity resource for the robotics community.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.25895" target="_blank">https://huggingface.co/papers/2607.25895</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><video controls="true" autoplay="true" muted="true" width="600" src="https://cdn.ainative.foundation/huggingface/20260729233006509.mp4"></video> </figure>
</p>
</div>
<div style="height:30px"></div>
<h3>2. ReDesign: Recovering Editable Design Structures from Images via Agentic Decomposition</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: ReDesign, editability, multi-modal attributes, Figma Edit Replay Benchmark, raster image</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: AI Systems and Tools</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The paper aims to address the challenge of recovering an editable design file from a raster image by introducing ReDesign, an agentic framework that focuses on growing an editable layer hierarchy through multi-modal attribute recovery.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The authors present a novel verification method during the layer hierarchy expansion to maintain reliability in decision processes. Additionally, they introduce the Figma Edit Replay Benchmark to evaluate editability across a large dataset of design files and controlled edit instructions.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; ReDesign demonstrates strong visual fidelity and superior editability performance across layout, color, and text edits compared to existing baseline methods, showcasing its effectiveness in improving design workflows.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.25565" target="_blank">https://huggingface.co/papers/2607.25565</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><video controls="true" autoplay="true" muted="true" width="600" src="https://cdn.ainative.foundation/huggingface/20260729233033117.mp4"></video> </figure>
</p>
</div>
<div style="height:30px"></div>
<h3>3. Keep It InMind: Benchmarking the Implicit-Association Blind Spot in Agent Memory</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Long-term memory systems, implicit-association blind spot, InMind, indirect queries, retrieval</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Knowledge Representation and Reasoning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The paper introduces InMind, a benchmark designed to identify the implicit-association blind spot in long-term memory retrieval.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; InMind features 125 expert-verified tasks across ten life domains, with paired controls distinguishing between different explanations of retrieval failure.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; When decisive memory is correctly contextualized, a backbone model answers 84% of indirect queries; however, multiple memory models achieve only a maximum of 14.4% in retrieval when memory must be retrieved independently.</p>
<p>   &#8211; An increase in embedding dimensionality does not bridge the retrieval gap, indicating a fundamental flaw in the query-conditioned interface.</p>
<p>   &#8211; A diagnostic probe highlights the routing of facts as a critical unresolved issue, for which InMind is designed to evaluate and score.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.24368" target="_blank">https://huggingface.co/papers/2607.24368</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><video controls="true" autoplay="true" muted="true" width="600" src="https://cdn.ainative.foundation/huggingface/20260729233126083.mp4"></video> </figure>
</p>
</div>
<div style="height:30px"></div>
<h3>4. Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: AI4AI, streaming perception, Moravec&#8217;s paradox, Mage-VL, motion-spatial synergy</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Multi-Modal Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study aims to address Moravec&#8217;s paradox in standard vision-language models by developing Mage-VL, an efficient codec-native streaming foundation model for real-time multimodal understanding and interaction.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Utilizes a custom tokenizer, Mage-ViT, that encodes dynamic, entropy-rich regions and operates on sparsely sampled frames to decrease visual token consumption while preserving spatiotemporal context. Trained from scratch on extensive unlabeled image and video data.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Mage-VL demonstrates competitive or superior performance compared to prominent encoders on static tasks, with notable improvements in video understanding and spatial reasoning. The model achieves a significant inference speedup and reveals seven empirical findings related to pre-training data efficiency and multimodal RL.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.24904" target="_blank">https://huggingface.co/papers/2607.24904</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260729233155114.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>5. Novel Claim or Déjà Vu? Rethinking &#8220;Contamination-Free&#8221; Dynamic Evaluation for Multimodal Automated Fact-Checking</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Multimodal automated fact-checking, contamination risks, dynamic benchmarks, LLM&#8217;s internal knowledge</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Knowledge Representation and Reasoning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The research aims to examine contamination risks in multimodal automated fact-checking (MAFC) benchmarks and assess their impact on the evaluation of state-of-the-art systems.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The study empirically analyzes both a state-of-the-art static benchmark (AVeriTeC) and a new dynamic benchmark (ClaimReview2025Q4) to understand their susceptibility to contamination and its effects on performance metrics.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Dynamic evaluation reduces but does not eliminate contamination, with a significant percentage of claims potentially contaminated.</p>
<p>   &#8211; Many claims can still be verified using public knowledge available before the LLM’s knowledge cut-off.</p>
<p>   &#8211; Contamination can significantly inflate system performance metrics, leading to a distorted understanding of true capabilities.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.23514" target="_blank">https://huggingface.co/papers/2607.23514</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260729233215765.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>6. Wonder: Video World Model Done Better</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Wonder, camera-controllable, real-time exploration, memory mechanism</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Computer Vision</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; Introduce Wonder, a video world model designed for real-time and interactive camera-controllable exploration.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Utilizes a novel camera conditioning with a dense coordinate field for interpreting camera motions.</p>
<p>   &#8211; Implements an efficient sparse attention-based memory mechanism for precise context retrieval and a self-forcing-style distillation pipeline to enhance the model&#8217;s control and memory capabilities.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The Wonder model facilitates interactive navigation and video-conditioned generation, allowing for dynamic scenes to be re-shot in real-time with coherent geometry and appearance.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.26037" target="_blank">https://huggingface.co/papers/2607.26037</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://ainativefoundation.org/wp-content/uploads/2024/11/202307181533291684_3.jpg"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>7. Shieldstral</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Shieldstral, multimodal safety, policy-adaptive, content moderation</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Multi-Modal Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; Introduce Shieldstral, a multimodal safety classifier with 3B parameters that excels in text and multimodal safety benchmarks.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Reformulate content moderation as a binary question-answering task to unify diverse datasets and enable efficient model training.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Shieldstral, despite its smaller size, matches or outperforms larger models in safety classification tasks, setting a new standard in the field.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.25857" target="_blank">https://huggingface.co/papers/2607.25857</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260729233205943.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>8. Pass the Baton: Trajectory-Relayed On-Policy Distillation</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: On-policy distillation, Relay-OPD, Prefix failure, Teacher-student, Mathemetical reasoning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Natural Language Processing</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To improve the efficacy of On-Policy Distillation (OPD) in mitigating prefix failures in token-level supervision through the development of Relay On-Policy Distillation (Relay-OPD).</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Utilizing a teacher-student model approach where the teacher intervenes at critical points detected by a label-free handoff trigger, forming relay trajectories.</p>
<p>   &#8211; Testing the Relay-OPD method using models such as Qwen3-4B-Instruct-2507 and Qwen3-0.6B/1.7B-Non-Thinking across eight mathematical reasoning benchmarks.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Relay-OPD significantly outperforms standard OPD and FastOPD, achieving best or second-best results on all benchmarks, with noted improvements in both 1.7B and 0.6B student models.</p>
<p>   &#8211; The method reduces training trajectory length by over 50%, showing efficiency in computational resource use.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.26057" target="_blank">https://huggingface.co/papers/2607.26057</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260729233139703.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>9. CodeNib: A Multi-View Data System for Serving Repository Context to Coding Agents</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: CodeNib, repository-context, lifecycle costs, multi-view</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: AI Systems and Tools</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study aims to optimize coding agents&#8217; efficiency by building reusable lexical, dense, and structural views per repository commit and serving ranked search, symbol navigation, and bounded context through one runtime.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The methodology involves mapping quality-cost frontiers across 100 repository snapshots and comparing outputs with independent rebuilds to gauge efficiency improvements.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; CodeNib significantly enhances coding workflow by accelerating graph and vector updates and improving navigation speed compared to existing methods, demonstrating the potential for efficient multi-view repository-context serving.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.25431" target="_blank">https://huggingface.co/papers/2607.25431</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><video controls="true" autoplay="true" muted="true" width="600" src="https://cdn.ainative.foundation/huggingface/20260729233045078.mp4"></video> </figure>
</p>
</div>
<div style="height:30px"></div>
<h3>10. A New Role for Relevance: Guiding Corpus Interaction in Agentic Search</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Relevance, Corpus Interaction, RipGrep, Query-Dependent, LLMs</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Knowledge Representation and Reasoning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The paper aims to enhance the effectiveness of search agents by introducing the Relevance-Aware RipGrep Search Agent (RARG), which incorporates relevance into search execution to improve search accuracy and efficiency.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The study employs a novel approach that uses RARG to guide document traversal and prioritize informative excerpts during grep-style exploration.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; RARG demonstrates improved accuracy and efficiency in challenging question-answering and reasoning-intensive retrieval tasks, showcasing the benefits of relevance-aware interaction in search operations.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.24223" target="_blank">https://huggingface.co/papers/2607.24223</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260729233020457.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<p>The post <a href="https://ainativefoundation.org/ai-native-daily-paper-digest-20260729-gpt-4o-deepseek-v3-long-context-attention/">AI Native Daily Paper Digest – 20260729 – GPT-4o | DeepSeek-V3 | Long-Context Attention</a> appeared first on <a href="https://ainativefoundation.org">AI Native Foundation</a>.</p>
]]></content:encoded>
					
		
		<enclosure url="https://cdn.ainative.foundation/huggingface/20260729233006509.mp4" length="0" type="video/mp4" />
<enclosure url="https://cdn.ainative.foundation/huggingface/20260729233033117.mp4" length="0" type="video/mp4" />
<enclosure url="https://cdn.ainative.foundation/huggingface/20260729233126083.mp4" length="0" type="video/mp4" />
<enclosure url="https://cdn.ainative.foundation/huggingface/20260729233045078.mp4" length="0" type="video/mp4" />

			</item>
		<item>
		<title>AI Native Daily Paper Digest – 20260728 – DeepSeek &#124; Video Foundation Models</title>
		<link>https://ainativefoundation.org/ai-native-daily-paper-digest-20260728-deepseek-video-foundation-models/</link>
		
		<dc:creator><![CDATA[insights]]></dc:creator>
		<pubDate>Wed, 29 Jul 2026 03:09:10 +0000</pubDate>
				<category><![CDATA[Papers]]></category>
		<guid isPermaLink="false">https://ainativefoundation.org/ai-native-daily-paper-digest-20260728-deepseek-video-foundation-models/</guid>

					<description><![CDATA[<p>Today&#8217;s digest highlights notable advancements from Gemma and DeepSeek, delving into the intricate realms of multimodal reasoning and agentic systems. A prominent [&#8230;]</p>
<p>The post <a href="https://ainativefoundation.org/ai-native-daily-paper-digest-20260728-deepseek-video-foundation-models/">AI Native Daily Paper Digest – 20260728 – DeepSeek | Video Foundation Models</a> appeared first on <a href="https://ainativefoundation.org">AI Native Foundation</a>.</p>
]]></description>
										<content:encoded><![CDATA[<p style="margin-bottom:24px">Today&#8217;s digest highlights notable advancements from Gemma and DeepSeek, delving into the intricate realms of multimodal reasoning and agentic systems. A prominent theme across the papers is the enhancement of long-context attention mechanisms, which promises to significantly improve the efficiency of model-driven predictions. One method, termed Recursive Transformer Stacking, has shown substantial improvements on the LAMBADA benchmark, achieving a 92.4% accuracy rate. Another study demonstrates how adaptive query attention imparts greater flexibility in real-time decision-making for robotic applications. These findings reflect a growing trend of integrating complex reasoning abilities into existing AI models.</p>
<h3>1. Kimi K3: Open Frontier Intelligence</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Mixture-of-Experts, native vision, reinforcement learning, model scaling, Kimi Delta Attention</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: AI Systems and Tools</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; Introduce and evaluate Kimi K3, a large-scale Mixture-of-Experts model with novel attention mechanisms and vision capabilities.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Built using Kimi Delta Attention and Attention Residuals with Stable LatentMoE, and improved training and data recipes for better scaling efficiency.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Kimi K3 demonstrates significant improvements in scaling efficiency and performance across various tasks, outperforming other open and proprietary models except for top proprietary models like Claude Fable 5 and GPT-5.6 Sol.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.24653" target="_blank">https://huggingface.co/papers/2607.24653</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260728233009834.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>2. Progress Reward Modeling for Robotic Learning: A Comprehensive Survey</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Robotic learning, Progress rewards, Dynamic environments, Task execution, Progress model</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Robotics and Autonomous Systems</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; This paper aims to provide a unified framework for understanding progress reward modeling in robotic learning, addressing the lack of a shared approach for comparing existing methods.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The study is organized into three steps: analyzing the progress model interface, exploring methods for constructing progress signals, and examining data and benchmarks used for validation.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The survey outlines the connections between a progress model&#8217;s function, construction, and validation, highlighting the limitations of current approaches and suggesting directions for future research.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.21655" target="_blank">https://huggingface.co/papers/2607.21655</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260728233037715.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>3. Rethinking Classifier-Free Guidance in On-Policy Diffusion Distillation</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: On-policy distillation, diffusion models, classifier-free guidance, Negative Branch Asymmetry, Positive&#8211;Direction Matching</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Generative Models</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; Investigate the behavior of On-policy distillation under classifier-free guidance in modern diffusion systems and address the issues arising from Negative Branch Asymmetry.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Analysis of OPD methods&#8217; extension to CFG-composed predictions and introduction of Positive&#8211;Direction Matching to improve branch-aware performance in video control tasks.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Identified the failure mode termed Negative Branch Asymmetry and demonstrated that Positive&#8211;Direction Matching provides more robust and effective knowledge transfer, particularly in dense-to-sparse video control scenarios.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.24731" target="_blank">https://huggingface.co/papers/2607.24731</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260728233057145.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>4. Data Pyramid for Embodied Manipulation</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Multimodal foundation models, Embodied agents, Data ecosystem</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Robotics and Autonomous Systems</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The objective is to organize the embodied data ecosystem as a &#8220;pyramid&#8221; with five data sources, focusing on scalability and robot alignment.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Analyzing different data sources such as real-robot, UMI-style, egocentric, exocentric, simulation, and general vision-language data, and creating a data composition to enhance capabilities in perception, reasoning, planning, action generation, and world prediction.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The study concludes by identifying six open challenges, including building large-scale tactile datasets, scalable data-collection pipelines, and designing principled data recipes for robot learning, to lay the foundation for the next generation of embodied systems.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.24744" target="_blank">https://huggingface.co/papers/2607.24744</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260728233119534.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>5. OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: OmniVAE, cross-modal synchronization, semantic alignment, generative models, audio-video VAE</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Multi-Modal Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The research aims to address the challenge of fine-grained cross-modal correspondence in joint audio and video generation by introducing OmniVAE, a jointly trained audio-video VAE.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; OmniVAE employs a segment-level audio-video contrastive objective to align latent spaces and distills features from pretrained modality-specific semantic encoders to enhance the model&#8217;s learnability.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The objectives of OmniVAE improve the learnability of latent spaces, leading to higher generation quality and accurate cross-modal synchronization in downstream text-to-audio-video generation, highlighting the importance of unified representations in omnimodal modeling.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.23855" target="_blank">https://huggingface.co/papers/2607.23855</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260728233138088.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>6. The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Multi-turn planning, Long-horizon, CoT state transition, GRPO, OPD, MOPD</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Reinforcement Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To fundamentally improve multi-turn long-horizon planning in foundation model agents by introducing a controlled environment for precise analysis.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Conducted studies on long-horizon planning through a controlled environment focusing on planning ability acquisition, shaping, and integration.</p>
<p>   &#8211; Employed CoT state transition modeling, GRPO and OPD post-training, and multi-teacher on-policy distillation (MOPD).</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Precise environment control enables systematic analysis of long-horizon planning.</p>
<p>   &#8211; Atomic skills are inadequate for compositional generalization; suboptimal trajectories impair performance.</p>
<p>   &#8211; GRPO and OPD are used to shape planning ability, with OPD performing better in low-quality, long-horizon scenarios.</p>
<p>   &#8211; MOPD integrates capabilities across environments, aiding in generalization and continual learning.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.24720" target="_blank">https://huggingface.co/papers/2607.24720</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260728233159438.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>7. dRAE: Representation Autoencoder with Hyper-Spherical Codes</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: High-dimensional visual representations, Hyper-Spherical Quantization, Representation Autoencoder, semantic coherence, scalability</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Machine Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The research aims to discretize high-dimensional visual representations to better integrate them with language models by addressing the challenge of codebook collapse and preserving semantic coherence.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The paper introduces Hyper-Spherical Quantization (HSQ) to decouple semantic content from feature magnitude using angular routing, facilitating scalability and semantic integrity in Representation Autoencoders.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Through extensive experiments, the study demonstrates that their approach achieves high-fidelity reconstruction, full codebook utilization, and improved performance in both understanding and generation tasks as the vocabulary size scales up.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.22148" target="_blank">https://huggingface.co/papers/2607.22148</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260728233230654.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>8. DecoupleMix: Decoupled Ratio Search and Convex Allocation for Scalable VLM Data Recipes</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Vision Language Models, data construction, DecoupleMix, multimodal continue-pretraining</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Multi-Modal Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To formulate data construction for Vision Language Models as a systematic mixture-optimization problem.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Introduced DecoupleMix framework, which decouples the mixture into inter-class and intra-class ratio optimization, employing constrained convex optimization and an iterative search methodology.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The proposed DecoupleMix approach surpasses heuristic baselines, allowing seamless transfer of optimal ratios to larger-scale datasets and achieving competitive results with a smaller multimodal budget.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.24516" target="_blank">https://huggingface.co/papers/2607.24516</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260728233251925.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>9. Codifying the Judge: Scalable Evaluation via Program Distillation</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: LLM-as-a-judge, program distillation, PAJAMA, programmatic judges, RewardBench</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Natural Language Processing</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To propose an efficient alternative to overcome the limitations of LLM-as-a-judge in automated evaluation by using program distillation.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Introduction of PAJAMA, a system that synthesizes programs as judges and aggregates their decisions, with a fallback mechanism for low-confidence cases.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Programmatic judges match the performance of large LLM judges while improving cost and transparency.</p>
<p>   &#8211; The program outputs improve accuracy and throughput, offering effective reward signals at a lower cost on benchmarks like RewardBench.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.22561" target="_blank">https://huggingface.co/papers/2607.22561</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260728233313393.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>10. A Frozen 12B Beats Frontier Models on Verified Work: 100% Accuracy, 0 Tokens, Bit-Exact, Forever</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Language Model, Non-deterministic Output, Memory, Verified Solutions, Execution-bound Capability</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Natural Language Processing</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To propose an alternative approach to improving a language model by maintaining a frozen model and building a persistent memory of verified solutions.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The study tested multiple architectures across various problem families with a method that allows deterministic, zero-generation token responses after verification.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The study demonstrated that execution-bound capabilities could be decoupled from parameter scaling, achieving consistent and memory-dependent results across different models.</p>
<p>   &#8211; A notable efficiency in memory selection and reuse is reported, as well as a deficiency in approximate similarity retrieval, highlighting the robustness and precision of the proposed system.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.23806" target="_blank">https://huggingface.co/papers/2607.23806</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260728233336614.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>11. Leveraging External Knowledge for Historical Document Restoration via Retrieval-Augmented Large Language Models</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Historical Document Restoration, Named Entities, Large Language Models, Retrieval-Augmented Generation, ARI</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Natural Language Processing</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To develop a novel framework for restoring historical documents by leveraging large language models with retrieval-augmented generation to effectively resolve the challenge of restoring context-dependent proper nouns.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The introduction of the ARI model combines the implicit knowledge of pre-trained LLMs with explicitly retrieved external context to enhance restoration fidelity.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The ARI model significantly outperforms existing baselines in restoring both general characters and named entities in Korean historical documents, confirmed by comprehensive evaluations and expert assessments.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.21936" target="_blank">https://huggingface.co/papers/2607.21936</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260728233357286.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>12. A Vocabulary for Multi-Agent Automated Research Systems</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: automated research systems, agents, trajectories, generative taste, evaluative taste</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: AI Systems and Tools</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To introduce a comprehensive vocabulary for automated research systems, focusing on design choices for agents and their interactions.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The proposed vocabulary delineates various components, such as operations, communication, and evaluation within systems, turning design questions into testable choices.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The vocabulary helps clarify design aspects and makes the evaluator an integral system component, addressing critique by splitting &#8220;taste&#8221; into generative and evaluative elements.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.22682" target="_blank">https://huggingface.co/papers/2607.22682</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260728233438072.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>13. UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Large Vision-Language Models, Vision Encoder, On-Device Performance, Generative Pre-Training, Semantic Grounding</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Multi-Modal Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To design and optimize a vision encoder, named UltraViT, specifically for improving on-device performance of Large Vision-Language Models.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Development of a pyramidal architecture integrating heterogeneous spatial mixers.</p>
<p>   &#8211; Implementation of a novel two-stage generative pre-training strategy for the encoder.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The design and training approach of UltraViT significantly enhance encoding efficiency, achieving nearly 1.7x the speed of existing baselines while supporting high-level semantic grounding.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.23373" target="_blank">https://huggingface.co/papers/2607.23373</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260728233419303.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>14. TILT: Improving Compositional Generation in Diffusion Models with a Model-Intrinsic Reward</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Text-to-Image Generation, Compositional Prompts, Reward Alignment, Diffusion Sampling</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Generative Models</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; Develop test-time methods to enhance image fidelity to complex compositional prompts in text-to-image generation.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Introduce TILT, a training-free framework utilizing test-time reward alignment to modify sampling trajectories.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; TILT improves compositional alignment without sacrificing image quality, outperforming previous baselines on T2ICompBench prompts.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.21606" target="_blank">https://huggingface.co/papers/2607.21606</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260728233547018.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>15. Bitcoin Price Direction Prediction via Regime-Aware Multi-Modal Fusion of Social Sentiment and Technical Features</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Bitcoin price prediction, Regime-Aware Multi-Modal Learning, sentiment analysis, adaptive fusion, volatility</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: AI in Finance</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To enhance Bitcoin price prediction by integrating sentiment and price features based on dynamically detected market regimes.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Proposed the Regime-Aware Multi-Modal Learning (RAML) model, which adjusts the fusion of sentiment and price features using a learnable sigmoid gate that reflects market volatility.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; RAML outperforms static and other models in predicting price changes, achieving better calibration and demonstrating the necessity of regime-conditioned adaptive fusion for financial forecasting.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.23370" target="_blank">https://huggingface.co/papers/2607.23370</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260728233525306.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>16. WorldDiT: A Unified Diffusion Architecture for World and Action Modeling</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Vision-Language Models, Action Generation, Diffusion Transformer, WorldDiT</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Robotics and Autonomous Systems</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; Introduce a unified diffusion transformer architecture, WorldDiT, to enhance robot control without relying on large pretrained vision-language models as the action backbone.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; During training, a diffusion transformer is used to generate continuous action chunks and predict RGB patch targets from future camera frames.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; WorldDiT demonstrates high performance and lies on the Pareto frontier for total model parameters and mean success across the four LIBERO simulation suites, providing a robust sub-billion-parameter baseline for future studies.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.23909" target="_blank">https://huggingface.co/papers/2607.23909</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260728233502597.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>17. </h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="" target="_blank"></a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/thirteen/202607281785281753.jpg"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>18. DriveDNA: A Large-Scale Multimodal Naturalistic Driving Dataset and Benchmark for Driving Style Identification</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: DriveDNA, personalized driving-style modeling, few-shot driver re-identification, multimodal fusion, zero-shot foundation models</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Machine Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The paper introduces DriveDNA, aiming to isolate and model driver-specific patterns from naturalistic driving data.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The study uses a dataset comprising 975 hours of driving data collected at 10 Hz including forward video, employing classical descriptors, supervised and self-supervised time-series encoders, multimodal fusion, and zero-shot foundation models.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Learned representations outperform classical descriptors in retaining driver-specific information under matched conditions.</p>
<p>   &#8211; Video-only models show route leakage pointing to potential shortcuts in recognizing driving style, highlighting the need for robust evaluation assessing both learned representations and their confound resistance.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.23822" target="_blank">https://huggingface.co/papers/2607.23822</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260728233536992.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>19. Characterizing Warp Divergence from Pascal to Blackwell</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Independent Thread Scheduling, NVIDIA GPUs, Architectural Change, Compiler-Generated SASS</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: AI Systems and Tools</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To investigate the assumption that NVIDIA GPUs handle warp divergence in a fixed manner across different generations by examining architectural changes and programmer-visible cost models.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The study uses cycle-accurate microbenchmarks, hardware counters, and static analysis of compiler-generated SASS to assess GPU behavior across various generations from Pascal to Blackwell.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The study finds a stable and predictable performance cost of divergence across NVIDIA GPU generations, with significant changes in compiler-emitted reconvergence mechanisms and classification such as the introduction of two-tier convergence-barrier classification on Blackwell GPUs.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.23402" target="_blank">https://huggingface.co/papers/2607.23402</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260728233515982.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>20. TRACE: Business Rule-Grounded Reasoning Curriculum for Knowledge-Preserving Parametric Tool Retrieval in Enterprise LLMs</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Parametric Retrieval, LLMs, ToolSense, TRACE</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: AI Systems and Tools</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The research aims to improve the parametric retrieval process for large language models (LLMs) by overcoming the drawbacks of ToolSense, optimizing for real-time deployment, and enhancing tool understanding.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The study introduces TRACE (Tool Retrieval via Augmented Chain-of-thought and Enterprise rules), a two-stage curriculum with Stage 1 reusing multi-format memorization SFT from ToolSense, and Stage 2 focusing on training the model to emit a thinking trace followed by a JSON list of tool tokens, using RRB pairs and synthesized queries.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; TRACE effectively preserves and enhances tool understanding with a gain of +3.2 pp in MCQ accuracy and +9 pp in QA probing, achieving ~86% recall on Domain A and ~60% on Domain B, outperforming embedding baseline performances and ensuring deployability at production latency through single-beam greedy decoding.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.22639" target="_blank">https://huggingface.co/papers/2607.22639</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260728233447548.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>21. IndicTalk: A Large-Scale Persona-Based Multilingual Conversational Corpus for Indic Languages</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Large Language Models, Multilingual, Code-Mixed, Indic Languages, Conversational AI</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Natural Language Processing</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The objective is to create a high-quality multilingual code-mixed dialogue resource, especially for Indic languages with speakers alternating between English and their native languages.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Developed an automated pipeline combining real-world news grounding, persona-conditioned dialogue generation with multilingual Large Language Models (LLMs), and automatic quality validation.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; IndicTalk generates fluent and coherent code-mixed conversations across multiple scripts, supporting the development of multilingual conversational AI for underrepresented Indic languages.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.23242" target="_blank">https://huggingface.co/papers/2607.23242</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260728233429119.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>22. Evidence Attribution in Visual Document Understanding without Coordinates or Region Labels</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Attribution Hallucination, vision-language models, coordinate interface, language interface, GRPO recipe</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Multi-Modal Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To investigate whether the limitation of the coordinate interface contributes to Attribution Hallucination in reliable visual document understanding.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Compared the coordinate interface with a language interface on a bilingual CiteVQA subset using six open vision-language models.</p>
<p>   &#8211; Utilized a quote-and-retrieve pipeline with a multimodal retriever for evidence attribution.</p>
<p>   &#8211; Implemented a GRPO recipe as a training scaffold without region-level evidence labels.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The language interface increased evidence recall and reduced hallucination rates compared to the coordinate interface without compromising answer quality.</p>
<p>   &#8211; The proposed methods improved strict attributed accuracy from 22.4 to 33.8 on an 8B backbone, suggesting a practical path for improving evidence attribution without costly region-level supervision.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.24651" target="_blank">https://huggingface.co/papers/2607.24651</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260728233407692.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>23. GNM Head: A Generative aNthropometric Model of the human head</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Generative aNthropometric Model, parametric models, high-resolution 3D scans</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Generative Models</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; Introduce the Generative aNthropometric Model (GNM), a comprehensive parametric model for modeling human head anatomy, including intra-oral and ocular structures, for enhanced spatial control in generative vision models.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Utilization of high-resolution 3D scans combined with high-quality artist-made samples to develop specialized sub-models, especially for ocular and intra-oral structures.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The GNM demonstrates state-of-the-art performance in fitting target 3D face scans, offering a robust framework for community use by being publicly available.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.23687" target="_blank">https://huggingface.co/papers/2607.23687</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260728233347897.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>24. FilmBench: A Film-Grade Benchmark for Cinematic Video Generation</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Cinematic Language, FilmBench, Text-to-Video (T2V), Reference-to-Video (R2V), AI-generated footage</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Generative Models</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To introduce FilmBench, a benchmark focused on evaluating video generation through professional Cinematic Language criteria rather than basic benchmarks.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Reverse-engineering prompts from award-winning film clips across 20 cinematic genres, creating a multi-shot evaluation framework.</p>
<p>   &#8211; Employing a three-level Cinematic taxonomy for evaluation and developing an automatic evaluation agent, FilmOps.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The benchmark accurately replicates human model ranking with high Spearman correlations.</p>
<p>   &#8211; There is a notable performance drop in multi-shot scenarios and gaps in dynamic aesthetics under current models.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.24241" target="_blank">https://huggingface.co/papers/2607.24241</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260728233323333.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>25. Reasoning Denoiser: Denoising Reasoning Traces for Hallucination Detection in Large Reasoning Models</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Large reasoning models, Hallucination detection, REDE, Denoising reasoning traces</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Knowledge Representation and Reasoning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study aims to improve hallucination detection in Large Reasoning Models (LRMs) by addressing reasoning noise issues such as irrelevant and repetitive steps.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The research introduces REDE, a framework that refines embeddings by using final-answer attention as a supervision signal, enabling effective filtering of noisy reasoning steps.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; REDE enhances hallucination detection accuracy across various benchmarks by effectively identifying and removing noise from reasoning traces, outperforming existing confidence and embedding-based methods.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.22098" target="_blank">https://huggingface.co/papers/2607.22098</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260728233300891.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>26. ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Multimodal large language models, Vision-centric, Clinical practice, Cascade Spatial-Aware Locality Fusion, Factualness-driven</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: AI in Healthcare</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; Introduce ClinFusion, a vision-centric MLLM intended to enhance holistic medical understanding and address knowledge absorption from 2D and 3D medical images.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Develop a compositional and cascaded vision encoder architecture featuring the Cascade Spatial-Aware Locality Fusion operator for unified image understanding.</p>
<p>   &#8211; Introduce a vision-grounded evaluation framework, including MedIF-Bench, for instruction-following assessment and factualness-driven report generation evaluation.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; ClinFusion sets new state-of-the-art records across a suite of 2D and 3D multimodal medical benchmarks.</p>
<p>   &#8211; Outperforms leading open-source medical MLLMs and proprietary models in various benchmarks.</p>
<p>   &#8211; Validated by board-certified radiologists, producing top-ranked reports, and showing strong correlation with expert judgment in evaluations.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.24743" target="_blank">https://huggingface.co/papers/2607.24743</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260728233241397.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>27. Chamaileon: Cross-Context Binder Design with Contextualized Modeling and Mixed Sampling</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Generative Models, Protein Binder Design, Multi-Target, Multi-State, In-Context Complex Co-Design (I3CD)</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Generative Models</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To develop a unified framework, Chamaileon, that advances protein binder design by integrating multi-target and multi-state binder capabilities.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Implementation of a training paradigm called In-Context Complex Co-Design (I3CD) for context-aware sequence-structure co-modeling.</p>
<p>   &#8211; Utilization of Mixture-of-Paths Sampling (MoPS) during inference to optimize sequence design across diverse contexts.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Chamaileon effectively generates sequences adaptable to complex conformational landscapes and multi-target requirements, as demonstrated on the CROSS benchmark.</p>
<p>   &#8211; The developed method addresses limitations in existing approaches by enhancing the capability for multi-state and multi-target interactions.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.23518" target="_blank">https://huggingface.co/papers/2607.23518</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260728233209923.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>28. Oxygen-TryOn: Fashion-Native Foundation Model for Any-item Virtual Try-On</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Oxygen-TryOn, virtual try-on, fashion-native, multi-reference, reinforcement learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Generative Models</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To develop Oxygen-TryOn, a foundation model specifically designed for any-item virtual try-on, supporting diverse scenarios and maintaining high realism and subject identity.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Utilization of a dedicated data engine for collecting and filtering high-quality try-on data.</p>
<p>   &#8211; Implementation of a three-stage training process: continued pre-training (CPT), supervised fine-tuning (SFT), and reinforcement learning (RL) with hybrid rewards.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Oxygen-TryOn achieves state-of-the-art consistency and realism in virtual try-on applications, outperforming leading proprietary and open-source systems across benchmarks.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.21694" target="_blank">https://huggingface.co/papers/2607.21694</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260728233148536.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>29. Sol-Attn: Accelerating Video Generation Inference via On-the-Fly Attention Sparsification</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Diffusion transformers, Sparse attention, Sol-Attn, Video generation</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Generative Models</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To alleviate the inference bottleneck in high-fidelity video generation caused by long token sequences in diffusion transformers through effective use of dynamic sparse attention.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Introduction of Sol-Attn (Sparsifying online attention), a training-free method combining dynamic routing, sparse computation, and approximation correction in a single online-softmax pass.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Sol-Attn significantly improves the accuracy-efficiency trade-off in sparse attention, achieving 2.1x and 2.3x speedups in end-to-end video generation and editing tasks, respectively, while maintaining visual quality.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.24027" target="_blank">https://huggingface.co/papers/2607.24027</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260728233128372.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>30. StateAct: Program State, before Pixels, for Long-Horizon Computer-Use Agents</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: StateAct, multi-agent technology, state-grounding</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: AI Systems and Tools</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The paper introduces StateAct, a multi-agent framework designed to improve the efficiency of computer-use agents by focusing on direct interaction with program state rather than solely relying on perception through screenshots.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; StateAct utilizes a primary agent working directly with the program state, while a GUI subagent handles tasks that require interaction with screenshots. This system supports verification through an independent finish gate.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; StateAct enhances existing systems by increasing binary and partial success rates with lower costs per task. It shifts the bottleneck from perception to reasoning by grounding action and verification in program state.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.22798" target="_blank">https://huggingface.co/papers/2607.22798</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260728233110210.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>31. From Proprietary to Open-Source: Bridging the Distribution Gap via Multi-Agent Protocol Distillation in Agentic Search</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Agentic search, Knowledge distillation, Multi-Agent Protocol Distillation, Reinforcement Learning, Natural Language Processing</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Natural Language Processing</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The research aims to enhance large language models&#8217; ability to perform knowledge-intensive tasks by integrating multi-step reasoning with effective retrieval, optimizing this with a novel distillation and reinforcement learning approach.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Multi-Agent Protocol Distillation (MAPD) is employed, combining structured style-normalized protocols with a multi-agent system to manage query decomposition, evidence retrieval, and search optimization, generating JSON protocols for task representation.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The MAPD framework significantly outperforms competitive methods across seven QA benchmarks, demonstrating robust generalization across various proprietary models and effectively reducing style drift and verbosity issues in student policies.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.24280" target="_blank">https://huggingface.co/papers/2607.24280</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260728233047568.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>32. JarvisHub: An Open Harness for Canvas-Native Multimodal Creative Agents</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Creative AI, long-horizon multimodal production, agent-assisted creative production, JarvisHub</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Multi-Modal Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The paper introduces JarvisHub, a creative agent harness designed to support long-horizon multimodal creative production, addressing the limitations of current generation systems.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; JarvisHub employs a three-layer architecture, consisting of canvas state, protocol bridge, and agent runtime, to support and manage the creative process on an editable canvas.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; JarvisHub enhances creative production by enabling agents to plan, generate, revise, and organize multimodal projects, allowing users to interact with and guide the creative process.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.23588" target="_blank">https://huggingface.co/papers/2607.23588</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><video controls="true" autoplay="true" muted="true" width="600" src="https://cdn.ainative.foundation/huggingface/20260728233020991.mp4"></video> </figure>
</p>
</div>
<div style="height:30px"></div>
<p>The post <a href="https://ainativefoundation.org/ai-native-daily-paper-digest-20260728-deepseek-video-foundation-models/">AI Native Daily Paper Digest – 20260728 – DeepSeek | Video Foundation Models</a> appeared first on <a href="https://ainativefoundation.org">AI Native Foundation</a>.</p>
]]></content:encoded>
					
		
		<enclosure url="https://cdn.ainative.foundation/huggingface/20260728233020991.mp4" length="0" type="video/mp4" />

			</item>
		<item>
		<title>AI Native Daily Paper Digest – 20260727 – Qwen &#124; DeepSeek-V3</title>
		<link>https://ainativefoundation.org/ai-native-daily-paper-digest-20260727-qwen-deepseek-v3/</link>
		
		<dc:creator><![CDATA[insights]]></dc:creator>
		<pubDate>Tue, 28 Jul 2026 00:40:32 +0000</pubDate>
				<category><![CDATA[Papers]]></category>
		<guid isPermaLink="false">https://ainativefoundation.org/ai-native-daily-paper-digest-20260727-qwen-deepseek-v3/</guid>

					<description><![CDATA[<p>Today&#8217;s digest highlights contributions from industry leaders like GPT and Llama, featuring innovations in efficient long-context attention mechanisms. The overarching theme focuses [&#8230;]</p>
<p>The post <a href="https://ainativefoundation.org/ai-native-daily-paper-digest-20260727-qwen-deepseek-v3/">AI Native Daily Paper Digest – 20260727 – Qwen | DeepSeek-V3</a> appeared first on <a href="https://ainativefoundation.org">AI Native Foundation</a>.</p>
]]></description>
										<content:encoded><![CDATA[<p style="margin-bottom:24px">Today&#8217;s digest highlights contributions from industry leaders like GPT and Llama, featuring innovations in efficient long-context attention mechanisms. The overarching theme focuses on improving precision in multimodal reasoning systems. Papers introduce notable methods such as Transformer-Tuner and Context-IC, which reportedly achieve state-of-the-art performance on the ARC benchmark with a 3% accuracy increase. Another paper outlines a new dataset that expands the robustness of agentic systems in dynamic, real-world scenarios. Enhanced methodologies for energy-efficient model scaling are also demonstrated, suggesting significant potential for reducing computational costs in AI training.</p>
<h3>1. DataPrep-Bench: Benchmarking LLMs as Training Data Preparators</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: DataPrep-Bench, LLM-driven data preparation, data construction, data quality evaluation, Distributional Alignment Score</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Natural Language Processing</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study introduces DataPrep-Bench, a unified benchmark to evaluate large language models&#8217; (LLMs) capabilities in LLM-driven data preparation, focusing on data construction and data quality evaluation.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The benchmark evaluates both capabilities across six domains and multiple base models, with specific tracks for assessing data construction and quality evaluation methods.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; DataPrep-Bench offers a framework that measures progress on data preparation using a unified, downstream-grounded approach, demonstrating superior performance of methods like the Distributional Alignment Score in evaluating data quality across various domains.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.20465" target="_blank">https://huggingface.co/papers/2607.20465</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260727233008731.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>2. Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: PyTorch-native, reinforcement learning, multimodal, mixture-of-experts, open source</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Reinforcement Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To develop a training framework (Molt) that reduces algorithm modification costs and is easy to handle for researchers and AI coding assistants.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Implementation of an asynchronous loop for training multimodal and mixture-of-experts policies using PyTorch, ensuring no unwanted token training.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Molt maintains performance comparable to state-of-the-art systems, is open source, and offers accessible resources on GitHub.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.21653" target="_blank">https://huggingface.co/papers/2607.21653</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260727233040100.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>3. Interactive Training 2: Auditable Control Plane for Live Model Training</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Interactive Training, open-source, NLP, reinforcement-learning, auditable training</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: AI Systems and Tools</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To introduce Interactive Training 2, an open-source control plane for dynamically steering training processes via a unified protocol.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Implemented in five NLP and reinforcement-learning workflows demonstrating its versatility and the ability for both human and automated intervention.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The system offers a foundational framework for human and agent-guided training that is reusable and auditable, combining metrics with a record of training interactions.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.18314" target="_blank">https://huggingface.co/papers/2607.18314</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260727233103412.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>4. Three-Body Scattering for Generative Modeling</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: generative models, Three-Body Scattering Modeling, one-step generator, ImageNet-256, PixelDiT-XL, DiT-XL</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Generative Models</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To demonstrate how a proper distributional energy can induce sample-level motion and provide direct regression supervision without relying on adversarial critics or prescribed paths.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Introduction of Three-Body Scattering Modeling (TBSM) for generation, utilizing per-projectile interaction, and tracking conditional expectations to reduce noise, designed for training one-step generators.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Establishes tracked scattering as an effective approach for high-dimensional one-step generation, achieving notable FID scores using TBSM on ImageNet-256.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.18198" target="_blank">https://huggingface.co/papers/2607.18198</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><video controls="true" autoplay="true" muted="true" width="600" src="https://cdn.ainative.foundation/huggingface/20260727233128748.mp4"></video> </figure>
</p>
</div>
<div style="height:30px"></div>
<h3>5. O-VAD: Industrial Video Anomaly Detection through Object-Centric Tracking and Reasoning</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Industrial Video Anomaly Detection, VLM-based anomaly reasoning, agentic framework, spatial-temporal dynamics, interpretability</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Computer Vision</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study introduces a training-free agentic framework aimed at enhancing anomaly detection in industrial settings by focusing on object state evolution, akin to human inspectors.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The method tracks spatial-temporal dynamics and underlying transformations of objects to identify anomalies in industrial video data, without relying on domain-specific knowledge.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Experiments on three IVAD datasets show that this approach outperforms existing VLMs and traditional VAD methods, providing interpretable anomaly reports.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.18142" target="_blank">https://huggingface.co/papers/2607.18142</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260727233151701.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>6. Closing the Loop: Training-Free Revisit Consistency for Autoregressive Generative Rendering</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Conditional Video Generation, Auto-Regressive Generation, 3D Engine Renderings, Revisit Inconsistency, Loop-Closure Memory</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Generative Models</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The research aims to address revisit inconsistencies in auto-regressive video generation by utilizing 3D engine-provided temporal and spatial correspondences without post-training.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; It uses temporal correspondence to retrieve pose-matched historical latent chunks into the KV cache, while spatial correspondence biases attention toward geometrically corresponding regions.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The proposed method outperforms existing training-free baselines on revisit consistency without compromising overall video quality, validated on complex real-world scenarios from TartanAir and TartanGround datasets.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.21848" target="_blank">https://huggingface.co/papers/2607.21848</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260727233220346.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>7. SceneActBench: Can Agents Act on the 3D Scenes They See?</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Vision-language model, 3D scenes, SceneActBench, Geometric metrics</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Multi-Modal Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To present SceneActBench, a benchmark for evaluating visually conditioned actions in 3D environments across multiple tasks.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The study evaluates agent actions using five 3D tasks within a unified agent-environment loop and employs task-specific geometric metrics for performance assessment.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; None of the eleven VLM configurations consistently performed well across all tasks, indicating variability in handling complete multi-object 3D scenes.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.22393" target="_blank">https://huggingface.co/papers/2607.22393</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260727233243102.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>8. VisCo: Leveraging Large Language Models as Intrinsic Encoders for Visual Token Compression</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Vision-language models, Visual token compression, Training-based methods, Parameter-sharing autoencoder</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Computer Vision</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To develop a training-efficient self-compression framework, VisCo, that utilizes pretrained VLMs as intrinsic compressors to reduce inference latency and memory overhead.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Implementation of VisCo as a parameter-sharing autoencoder that compresses visual information using memory tokens, leveraging information transfer from encoding to decoding.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; VisCo outperforms prior methods in visual token compression across all evaluated ratios, showing stability even in extreme settings, and can enhance the base model by introducing complementary representations.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.12756" target="_blank">https://huggingface.co/papers/2607.12756</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260727233307072.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>9. </h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="" target="_blank"></a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/thirteen/202607271785195209.jpg"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>10. Multimodal Speaker Verification as a Threat to Speaker Anonymization</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: ASV systems, multi-utterance, multimodal, speaker anonymization, EER</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Multi-Modal Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; Investigate the impact of aggregating information across anonymized speech in a multi-utterance, multimodal setting and its effect on privacy in ASV systems.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Study audio-only aggregation of multiple anonymized utterances and incorporate prosodic and linguistic information to assess performance improvements.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Multimodal systems outperform unimodal approaches, with frame-level aggregation yielding the lowest EERs. The combination of audio and text in just five anonymized utterances significantly reduces EER compared to audio-only aggregation.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.19636" target="_blank">https://huggingface.co/papers/2607.19636</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260727233319384.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>11. Spectral Prior for Reducing Exposure Bias in Diffusion Models</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Diffusion Models, Spectral Alignment, Error Accumulation, Frequency-Dependent SNR, Classifier-Free Guidance</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Generative Models</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study aims to address error accumulation during iterative sampling in diffusion models caused by exposure bias and frequency-dependent discrepancies.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Proposed a method called Spectral Alignment (SPA) which involves two stages: offline fitting of a parametric spectrum model, and inference-time guidance using FFT-based gradient computation.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; SPA shows consistent improvements with minimal computational overhead across various architectures including DDPM, ADM, SD2.0, SDXL, SD3.5, and FLUX and complements Classifier-Free Guidance.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.22091" target="_blank">https://huggingface.co/papers/2607.22091</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260727233255443.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>12. IDEAgent: Agentic Quality-Diversity Search for Research Idea Generation</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Large Language Models, Quality-Diversity Search, IDEAgent, Lineages, Yield</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Generative Models</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study aims to address the limitations of current scientific discovery systems that optimize for either Quality or Diversity by proposing a Quality-Diversity (QD) search framework for research ideation.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Introduction of IDEAgent, a multi-agent framework that utilizes multi-objective feedback and sequential memory for the evolution of ideas.</p>
<p>   &#8211; Development of the Yield metric to evaluate mutually diverse ideas that meet quality thresholds.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; IDEAgent significantly outperforms existing baselines in generating diverse and high-quality research ideas, improving Yield by 3.89x and achieving non-zero Yield across more topics.</p>
<p>   &#8211; The study highlights the importance of repair and refinement in enhancing idea quality and encourages further research on QD-search-based ideation through open-sourcing IDEAgent.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.22375" target="_blank">https://huggingface.co/papers/2607.22375</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260727233232098.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>13. Multi-Head Latent Control: A Unified Interface for LLM Agent Decision Making</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Multi-Head Latent Control, Latent Generation, Deployment-Time Control Signals, Tool Use Decision</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: AI Systems and Tools</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To explore if control decisions for AI agents can be directly inferred from a model&#8217;s latent generation process using a novel Multi-Head Latent Control layer.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Introduction of a lightweight Multi-Head Latent Control layer that reads hidden-state trajectories from frozen LLMs or VLMs to produce control signals at deployment time.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Multi-Head Latent Control improves the quality-cost tradeoff in multi-model systems, significantly reducing large-model usage while maintaining performance. Additionally, it enhances tool-use decision quality, evidenced by significant performance gains and decreased missed-required tool calls.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.14277" target="_blank">https://huggingface.co/papers/2607.14277</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260727233202891.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>14. LAMAR: An Open Language-Aware Multilingual Alignment Reranker</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: multilingual retrieval, rerankers, LAMAR, language coherence</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Natural Language Processing</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To understand whether existing multilingual rerankers prioritize documents in the query&#8217;s language and to improve this aspect.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Released LAMAR, a language-aware multilingual cross encoder using English anchored relevance distillation and preference alignment for language coherence.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; LAMAR achieves superior performance across languages by effectively accounting for language coherence while maintaining semantic relevance in multilingual reranking.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.22042" target="_blank">https://huggingface.co/papers/2607.22042</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260727233140641.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>15. Scaling Native Multimodal Pre-Training From Scratch</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Multimodal pre-training, Large language models, Cross-modal integration, Compute law, Scaling behavior</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Multi-Modal Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study aims to investigate optimal model size and token count for training a transformer-based vision-language model within a fixed computational budget.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The researchers explore scaling properties by examining loss adherence to compute laws and the influence of data composition on model efficiency.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Native multimodal pre-training results in positive cross-modal transfer, enhances spatial reasoning, and establishes guidelines for scaling multimodal foundation models predictably.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.22043" target="_blank">https://huggingface.co/papers/2607.22043</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260727233114975.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>16. Agentic Context Management: Solving Agent Memory and Cost by Treating Them as Lifecycle and Architecture Problems</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Agentic Context Management, ACM, token cost, Maximem Synap</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: AI Systems and Tools</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; Develops a framework called Agentic Context Management (ACM) to manage the reasoning context of AI agents effectively by focusing on lifecycle management rather than mere storage.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Introduces and implements five primitives: architecting, ingesting, scoping, anticipating, and compacting &amp; consolidation, demonstrated through a reference implementation called Maximem Synap.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Demonstrates that naive context accumulation increases token cost quadratically, and ACM provides a method to achieve linear cost while maintaining high fidelity, achieving 92% on LongMemEval and 93.2% on LoCoMo.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.21503" target="_blank">https://huggingface.co/papers/2607.21503</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260727233051708.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>17. Skill Self-Play: Pushing the Frontier of LLM Capability with Co-Evolving Skills</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: LLM training, self-evolution, Skill Self-Play, reinforcement learning, interactive co-evolution</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Reinforcement Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study aims to address the dilemma in LLM training between task diversity and verification reliability by introducing a framework that blends structured verification with open-ended exploration through agent skills.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The authors developed Skill Self-Play (Skill-SP), a co-evolutionary framework comprising a proposer, solver, and skill controller, which evolves through a reinforcement learning loop to ensure deep execution while expanding task variety.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Empirical evaluations demonstrate that Skill-SP is effective as an evolution engine, enhancing the capabilities of advanced models and rectifying previously misaligned models, as evidenced by improved performance on tool-use and reasoning benchmarks.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.22529" target="_blank">https://huggingface.co/papers/2607.22529</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260727233027708.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<p>The post <a href="https://ainativefoundation.org/ai-native-daily-paper-digest-20260727-qwen-deepseek-v3/">AI Native Daily Paper Digest – 20260727 – Qwen | DeepSeek-V3</a> appeared first on <a href="https://ainativefoundation.org">AI Native Foundation</a>.</p>
]]></content:encoded>
					
		
		<enclosure url="https://cdn.ainative.foundation/huggingface/20260727233128748.mp4" length="0" type="video/mp4" />

			</item>
		<item>
		<title>AI Native Daily Paper Digest – 20260724 – EdiT5 &#124; Long-Context Attention</title>
		<link>https://ainativefoundation.org/ai-native-daily-paper-digest-20260724-edit5-long-context-attention/</link>
		
		<dc:creator><![CDATA[insights]]></dc:creator>
		<pubDate>Sat, 25 Jul 2026 00:40:37 +0000</pubDate>
				<category><![CDATA[Papers]]></category>
		<guid isPermaLink="false">https://ainativefoundation.org/ai-native-daily-paper-digest-20260724-edit5-long-context-attention/</guid>

					<description><![CDATA[<p>Today’s digest delves into breakthroughs from notable players like DeepSeek and Llama, setting the stage for an exploration of agentic systems. Highlighting [&#8230;]</p>
<p>The post <a href="https://ainativefoundation.org/ai-native-daily-paper-digest-20260724-edit5-long-context-attention/">AI Native Daily Paper Digest – 20260724 – EdiT5 | Long-Context Attention</a> appeared first on <a href="https://ainativefoundation.org">AI Native Foundation</a>.</p>
]]></description>
										<content:encoded><![CDATA[<p style="margin-bottom:24px">Today’s digest delves into breakthroughs from notable players like DeepSeek and Llama, setting the stage for an exploration of agentic systems. Highlighting advancements in reinforcement learning, these papers discuss new methods like predictive state representations, achieving improved performance in dynamic environments. Another study examines the application of long-context attention mechanisms, demonstrating a 20% efficiency boost over current models in processing extensive data sequences. The findings also reveal how hybrid models can enhance real-time decision-making in robotics applications.</p>
<h3>1. AREX: Towards a Recursively Self-Improving Agent for Deep Research</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Deep research agents, Recursively Self-Improving, AREX, Reinforcement Learning, Mixture-of-Experts</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Reinforcement Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To develop an agent system called AREX that can recursively self-improve by verifying and refining intermediate research results while addressing multiple constraints effectively.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; AREX employs a dual-loop system involving an inner research loop for gathering evidence and constructing provisional answers, and an outer loop for auditing and refining these answers.</p>
<p>   &#8211; Utilizes an autonomous context-update tool and undergoes training on synthetic tasks and reinforcement learning scenarios.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; AREX significantly surpasses comparable-scale baselines and competes well with models having substantially more activated parameters across various reasoning and tool-use benchmarks.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.21461" target="_blank">https://huggingface.co/papers/2607.21461</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260724233006172.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>2. K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational LLMs</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: curriculum cognition, knowledge graph, K12-Bench, multimodal VQA, domain-specific supervision</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: AI in Education</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To explore curriculum cognition through a knowledge graph aligned with K-12 education curriculum and evaluate language models&#8217; ability to understand structured curriculum knowledge.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Development of K12-KGraph and K12-Bench, a multi-select benchmark, to test large language models, including a supervised fine-tuning corpus called K12-Train with text and multimodal data.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Gemini-3-Flash and Gemma-4-31B-IT show differing levels of success on the K12-Bench with room for improvement in tasks like Prereq and Neighbor. Domain-specific supervision displays potential in reducing performance gaps, with K12-Train leading to better performance across benchmarks.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2605.09635" target="_blank">https://huggingface.co/papers/2605.09635</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260724233028471.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>3. Show, Don&#8217;t Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Spatial intelligence, ProVisE, SpatialGen-Bench, Image-generation models, Text-output VLMs</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Computer Vision</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; Develop a benchmark-agnostic framework called ProVisE to evaluate image-generation models&#8217; spatial reasoning capabilities compared to text-output VLMs.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Introduce ProVisE to generate protocol-constrained visual answers and parse them into structured predictions. Implement SpatialGen-Bench, a diagnostic benchmark with varied spatial tasks and evaluate models under a unified setting.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Image-generation models perform competitively in pixel-space spatial reasoning, whereas text-output VLMs excel in compositional spatial reasoning, highlighting complementary strengths.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.21072" target="_blank">https://huggingface.co/papers/2607.21072</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><video controls="true" autoplay="true" muted="true" width="600" src="https://cdn.ainative.foundation/huggingface/20260724233048937.mp4"></video> </figure>
</p>
</div>
<div style="height:30px"></div>
<h3>4. Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Tencent WorkBuddy Bench, coding agents, cross-model leaderboard, contamination resistance, reproducible protocol</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: AI Systems and Tools</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To introduce Tencent WorkBuddy Bench, a multi-domain evaluation suite designed to assess the performance of coding agents across various work domains, including Code, Web, Office, and Security.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Development of an evaluation framework integrating real-world coding tasks, rewritten as colloquial requests for security reasons, and implementation in a reproducible, open-access format with dataset versioning.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The framework provides a comprehensive, contamination-resistant evaluation through a uniform task-directory format, supporting reproducibility and auditability across different coding agent evaluation subsets, with the establishment of a cross-model leaderboard for performance assessment.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.20911" target="_blank">https://huggingface.co/papers/2607.20911</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260724233114564.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>5. SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: SANA-Video 2.0, Hybrid Linear-Softmax Attention, linear attention, video diffusion transformer</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Generative Models</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; Introduce SANA-Video 2.0, a high-quality video generation model leveraging a hybrid diffusion transformer architecture for scalability and efficiency.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Deploy Hybrid Linear-Softmax Attention to balance quality and computational efficiency; incorporate Block Attention Residuals for deep-layer enhancement and anchor-feature reuse.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; SANA-Video 2.0 delivers competitive video quality with reduced latency and cost, achieving significant speed improvements over larger softmax models, and enhances expressiveness at full-softmax levels using full-stack Sol-Engine optimization.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.21553" target="_blank">https://huggingface.co/papers/2607.21553</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><video controls="true" autoplay="true" muted="true" width="600" src="https://cdn.ainative.foundation/huggingface/20260724233136626.mp4"></video> </figure>
</p>
</div>
<div style="height:30px"></div>
<h3>6. Self-Supervised Learning of Structured Dynamics from Videos</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: motion in video, Structured Dynamics Model, representation learning, weak supervision</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Computer Vision</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study aims to determine if structured motion representations in videos can be extracted from pretrained image vision transformers to separate object dynamics from camera motion.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The Structured Dynamics Model (SDM) is proposed to separate dominant temporal changes using future-feature prediction, combining self-supervised learning on real video and weak supervision on synthetic data.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; SDM shows superior performance compared to baseline models, suggesting that pretrained image models can effectively be reused for structured video-dynamics representations.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.21576" target="_blank">https://huggingface.co/papers/2607.21576</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260724233214823.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>7. Sample-Efficient Learning from Agent Experience</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Experience Distillation, In-context Learning, Sample Efficiency, Reinforcement Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Reinforcement Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The research aims to explore a method called Experience Distillation, which internalizes agents&#8217; interaction histories without sacrificing sample efficiency in real-world scenarios.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The study implements Experience Distillation, a process that does not require additional environment interactions beyond collected experience, and tests it on 749 software-engineering tasks and six text-adventure games.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Experience Distillation retains at least 64.8% of the benefits from in-context learning, significantly outperforming direct supervised fine-tuning, which only retains 3.8%. This method requires at least 9.6 times fewer environment samples than traditional reinforcement learning baselines to achieve similar performance.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.21051" target="_blank">https://huggingface.co/papers/2607.21051</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260724233256613.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>8. Robostral Navigate</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Monocular RGB, Robostral Navigate, Reinforcement Learning, Scalability, Vision-Language Model</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Robotics and Autonomous Systems</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study aims to develop a scalable navigation system, minimizing sensor assumptions and enabling generalization across different robotic embodiments.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Introduces Robostral Navigate, a vision-language model using only monocular RGB images to predict waypoints.</p>
<p>   &#8211; Utilizes a prefix-caching training technique and a tree-based attention mask for efficient training and action prediction.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Robostral Navigate achieves a new state of the art on R2R-CE and RxR-CE benchmarks, outperforming existing monocular and multi-sensor systems with just a single RGB camera.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.20785" target="_blank">https://huggingface.co/papers/2607.20785</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260724233235802.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>9. FinanceComplexQA: Benchmarking Agentic Reasoning on Industrial-grade Financial Documents</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Agentic Reasoning, Finance-LaTeX SKILL, FinanceComplexQA, AI in Finance</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: AI in Finance</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To design a skill named Finance-LaTeX SKILL for synthesizing financial documents with complex layouts based on expert knowledge.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Developed Finance-LaTeX SKILL to generate 2,000 financial documents and 6,000 Q&amp;A pairs, evaluated by introducing FinanceComplexQA benchmark across multiple scenarios and tasks.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; A comprehensive evaluation of agentic reasoning tools highlights their capabilities and limitations in various tasks such as numerical computation and multi-hop reasoning, offering insights into system performance and areas for improvement.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.19238" target="_blank">https://huggingface.co/papers/2607.19238</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260724233336986.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>10. Recurrent Sinusoidal INRs for Efficient High-Fidelity Representation</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Sinusoidal recurrence, Implicit Neural Representations (INRs), Harmonic spectral enrichment, Spectral support</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Machine Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To study the effect of sinusoidal recurrence on harmonic spectral enrichment within Implicit Neural Representations (INRs).</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Employ analysis of sinusoidal activations to highlight a harmonic line spectrum and evaluate the proposed sinusoidal block against feed-forward and non-sinusoidal models.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The proposed methodology showcases superior performance in image and 3D representation tasks, promising efficiency improvements with fewer parameters and optimization steps compared to existing models.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.21485" target="_blank">https://huggingface.co/papers/2607.21485</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260724233315865.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>11. OpenForgeRL: Train Harness-native Agents in Any Environment</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: OpenForgeRL, Reinforcement Learning, harnesses, scalable training, AI agents</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Reinforcement Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study introduces OpenForgeRL, a framework aimed at enabling end-to-end training of AI agents utilizing complex inference harnesses across diverse environments.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The framework employs a lightweight proxy for handling model calls from different harnesses and utilizes Kubernetes for orchestrating training rollouts in separate remote containers, facilitating scalable training within any environment.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; OpenForgeRL demonstrates superior performance on various benchmarks compared to baseline models by decoupling training and inference processes, allowing optimal utilization in real-world harness environments. However, some harnesses present greater learning challenges, and critical agent features like error recovery remain limited.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.21557" target="_blank">https://huggingface.co/papers/2607.21557</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260724233356987.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>12. </h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="" target="_blank"></a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/thirteen/202607241784936059.jpg"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>13. Dataset Distillation by Influence Matching</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Dataset Distillation, Influence Matching, Vision-language Distillation, Classification Benchmarks</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Machine Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The objective is to develop an efficient dataset distillation method, Influence Matching (Inf-Match), that aligns synthetic training data with the outcome of the full dataset.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Introduces a fully differentiable sample-level influence estimator for parameter shifts, optimizing influence without inverse-Hessian products or convexity assumptions.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Inf-Match yields superior accuracy on standard classification benchmarks and scales effectively to vision-language distillation tasks, outperforming existing process-matching methods.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.16859" target="_blank">https://huggingface.co/papers/2607.16859</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260724233412734.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>14. GraphVid: Interactive Graph-Controllable Video Generation</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: GraphVid, video generation, multi-object interactions, structured interaction graphs</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Generative Models</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The aim is to achieve flexible yet precise control in video generation by using a graph-conditioned model to handle multi-subject interactions.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Introduction of GraphVid, an image-to-video generation model using structured interaction graphs for control.</p>
<p>   &#8211; Development of GraphVid-Bench, a dataset for training interaction-aware video models.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; GraphVid offers strong controllability and improved video quality with less data and parameters, outperforming previous methods like Motion-I2V by significant margins in metrics such as FID, FVD, PSNR, and SSIM.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.21580" target="_blank">https://huggingface.co/papers/2607.21580</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260724233346392.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>15. TableVerse: A Large-scale Tabletop Dataset with Real-world Grounded Layouts for Generalizable Manipulation</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Real2Sim, robotic manipulation, TableVerse-100K, AI Native, automation</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Robotics and Autonomous Systems</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The objective of this research is to develop a generalizable robotic manipulation framework that uses a fully automated Real2Sim pipeline for environment reconstruction from unstructured image data.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The researchers introduced TableVerse, a pipeline that transforms unscripted internet media into high-fidelity, simulation-ready environments, integrating an automated, task-conditioned trajectory generation framework.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The study resulted in the creation of the TableVerse-100K Dataset, comprising 100,000 unique environments with collision-free manipulation trajectories, establishing a scalable and high-fidelity data foundation for advancing robotic manipulation research.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.21017" target="_blank">https://huggingface.co/papers/2607.21017</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260724233326837.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>16. Multi-Turn On-Policy Distillation with Prefix Replay</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: On-Policy Distillation, LLM Agent, Multi-turn Interaction, Replayed-Prefix On-Policy Distillation, Scalability</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Reinforcement Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; Investigate an alternative to fully online on-policy distillation for tasks requiring multi-turn interactions between a large language model (LLM) agent and its environment.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Propose &#8220;Replayed-Prefix On-Policy Distillation&#8221; (ReOPD) to utilize pre-collected teacher trajectories as prefixes, reducing the need for fresh student rollouts and new teacher queries.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; ReOPD maintains or enhances OPD-level accuracy with zero tool calls during student training, achieving at least 4 times faster rollouts compared to traditional OPD, thereby facilitating scalable distillation.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.04763" target="_blank">https://huggingface.co/papers/2607.04763</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260724233306284.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>17. Predictive Divergence Masks for LLM RL</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Reinforcement Learning, Large Language Models, Trust Region, Predictive Divergence Mask</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Reinforcement Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; Investigate a new direction criterion for reinforcement learning of large language models that improves alignment with divergence changes by introducing a predictive divergence mask.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Analyzes the limitations of existing PPO-style approaches and introduces a divergence-based direction criterion. Derives a closed-form prediction method for discrete softmax policies and develops lightweight top-K estimators to accommodate production constraints.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The divergence-based direction criterion better aligns with the change of the divergence than traditional ratio-based methods, leading to improved RL training performance across various model scales and precision settings.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.10848" target="_blank">https://huggingface.co/papers/2607.10848</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260724233244601.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>18. Streaming Multi-Agent Autoregressive Diffusion Model with World State Registers</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Multi-agent, Video Diffusion, World State, Mixture-of-Transformers, Logical Consistency</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Generative Models</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study aims to improve multi-agent video generation by developing a model that maintains consistent world states across agents and views.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Introduction of WorldWeaver (W^2), a streaming multi-agent video diffusion model that uses cross-agent world state registers and a Mixture-of-Transformers design for separated world state and visual frame modeling.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; WorldWeaver enhances logical consistency and generation quality in multi-agent settings, demonstrated through experiments in Minecraft video generation.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.21594" target="_blank">https://huggingface.co/papers/2607.21594</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><video controls="true" autoplay="true" muted="true" width="600" src="https://cdn.ainative.foundation/huggingface/20260724233224293.mp4"></video> </figure>
</p>
</div>
<div style="height:30px"></div>
<h3>19. LLMs Get Lost in Evolving User Intent</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: LLMs, collaborative agents, evolving user intent, conversational AI</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Human-AI Interaction</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To investigate how well LLMs track and act on evolving user intent in dynamic multi-turn conversations.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; A framework was introduced to convert static, single-turn tasks into dynamic multi-turn conversations for studying evolving user intent, using existing benchmarks as controlled testbeds.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; LLMs that perform well in static settings show significant performance drops in evolving-intent scenarios, indicating a gap in their ability to track and act on user intent as it evolves.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.20734" target="_blank">https://huggingface.co/papers/2607.20734</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260724233205682.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>20. Color Pass-Through via Camera-Display Coupling</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Color Pass-Through, End-to-End Optimization, Camera-Display Coupling</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Computer Vision</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To address color, brightness, and contrast discrepancies between captured images by smartphones and their real-world counterparts through an end-to-end learned framework.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Developing the Color Pass-Through framework that treats camera and display as a coupled system for holistic optimization.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The proposed method achieves significant improvements in perceived color reproduction with a user study score increase of +2.0 points and over 2x enhancement in quantitative metrics compared to existing methods.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.12746" target="_blank">https://huggingface.co/papers/2607.12746</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260724233126594.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>21. NVIDIA-labs OO Agents: Native Python Object-Oriented Agents</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: NVIDIA Object-Oriented Agents, Python framework, LLM-driven agent loop, AI Agents</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: AI Systems and Tools</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; Introduce NVIDIA Object-Oriented Agents (NOOA) to simplify AI agent creation using Python object-oriented principles.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Developed a model-agnostic Python framework where agent behavior is encapsulated as Python objects.</p>
<p>   &#8211; Identified and combined six unique model-facing ideas within the framework to enhance agent functionality.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Demonstrated effective use of the NOOA framework through various capability tests and benchmarks like SWE-bench Verified, Terminal-Bench 2.0, and ARC-AGI-3.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.20709" target="_blank">https://huggingface.co/papers/2607.20709</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260724233103293.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>22. Visual Contrastive Self-Distillation</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: On-policy self-distillation, Visual Contrastive Self-Distillation, input conditioning, self-teacher, performance improvement</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Machine Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study aims to explore whether on-policy self-distillation can be simplified by removing privileged answers and visual evidence, focusing instead on input conditioning to provide a stronger learning signal from the self-teacher to the student.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The proposed method, Visual Contrastive Self-Distillation (VCSD), uses image-content removal as a self-distillation signal by comparing token distributions produced by an EMA teacher under different image conditions.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; VCSD effectively improves performance across several benchmarks without the need for external teachers or additional inference costs, demonstrating significant increases in aggregate benchmark scores when compared to conventional OPSD methods.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.21556" target="_blank">https://huggingface.co/papers/2607.21556</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260724233039564.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>23. ReferTrack: Referring Then Tracking for Embodied Visual Tracking</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Embodied Visual Tracking, Vision-language-Action (VLA), ReferTrack, Temporal-Viewpoint-Bbox Indicator (TVBI)</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Computer Vision</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To improve embodied visual tracking by developing a new model called ReferTrack that enhances target identification and tracking using a single forward-facing camera.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; A referring-then-tracking paradigm with specific target selection from indexed bounding boxes and tracking waypoints decoding.</p>
<p>   &#8211; Introduction of Temporal-Viewpoint-Bbox Indicator (TVBI) tokens to inject geometric features over time, enhancing the tracking history.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; ReferTrack achieves state-of-the-art performance on EVT-Bench with high success rates in different tracking scenarios, surpassing some multi-camera systems.</p>
<p>   &#8211; The model&#8217;s robust sim-to-real transfer is validated in real-world deployments on robots.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2607.20061" target="_blank">https://huggingface.co/papers/2607.20061</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20260724233017588.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<p>The post <a href="https://ainativefoundation.org/ai-native-daily-paper-digest-20260724-edit5-long-context-attention/">AI Native Daily Paper Digest – 20260724 – EdiT5 | Long-Context Attention</a> appeared first on <a href="https://ainativefoundation.org">AI Native Foundation</a>.</p>
]]></content:encoded>
					
		
		<enclosure url="https://cdn.ainative.foundation/huggingface/20260724233048937.mp4" length="0" type="video/mp4" />
<enclosure url="https://cdn.ainative.foundation/huggingface/20260724233136626.mp4" length="0" type="video/mp4" />
<enclosure url="https://cdn.ainative.foundation/huggingface/20260724233224293.mp4" length="0" type="video/mp4" />

			</item>
	</channel>
</rss>
