<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>AI Native Foundation</title>
	<atom:link href="https://ainativefoundation.org/feed/" rel="self" type="application/rss+xml" />
	<link>https://ainativefoundation.org/</link>
	<description></description>
	<lastBuildDate>Sat, 10 Oct 2026 04:04:32 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.3</generator>

<image>
	<url>https://ainativefoundation.org/wp-content/uploads/2024/05/cropped-favicon-32x32.png</url>
	<title>AI Native Foundation</title>
	<link>https://ainativefoundation.org/</link>
	<width>32</width>
	<height>32</height>
</image> 
	<item>
		<title>Global AI Native Industry Insights &#8211; 20261010 &#8211;  Anthropic &#124; Google &#124; OpenAI &#124; more</title>
		<link>https://ainativefoundation.org/global-ai-native-industry-insights-20261010-anthropic-google-openai-more/</link>
		
		<dc:creator><![CDATA[AINF]]></dc:creator>
		<pubDate>Sat, 10 Oct 2026 04:04:32 +0000</pubDate>
				<category><![CDATA[Global Industry]]></category>
		<guid isPermaLink="false">https://ainativefoundation.org/?p=11105</guid>

					<description><![CDATA[<p>Today's digest brings exciting announcements from Anthropic, Google, and OpenAI. The overarching theme is the expansion and integration of AI into new domains, showcasing advancements in both terrestrial and space-based applications. Anthropic has introduced Claude Dashboards and Claude Motion in beta, aiming to enhance AI interaction and functionality. Meanwhile, Google is venturing into the cosmos with the launch of its first test satellite outfitted with TPUs, marking a significant step towards establishing a space-based AI infrastructure. Concurrently, OpenAI is enhancing speed across its platforms by rolling out the Ultrafast Mode for GPT-6.1 Sol on API, Codex, and ChatGPT.</p>
<p>The post <a href="https://ainativefoundation.org/global-ai-native-industry-insights-20261010-anthropic-google-openai-more/">Global AI Native Industry Insights &#8211; 20261010 &#8211;  Anthropic | Google | OpenAI | more</a> appeared first on <a href="https://ainativefoundation.org">AI Native Foundation</a>.</p>
]]></description>
										<content:encoded><![CDATA[<p>Today&#8217;s digest brings exciting announcements from Anthropic, Google, and OpenAI. The overarching theme is the expansion and integration of AI into new domains, showcasing advancements in both terrestrial and space-based applications. Anthropic has introduced Claude Dashboards and Claude Motion in beta, aiming to enhance AI interaction and functionality. Meanwhile, Google is venturing into the cosmos with the launch of its first test satellite outfitted with TPUs, marking a significant step towards establishing a space-based AI infrastructure. Concurrently, OpenAI is enhancing speed across its platforms by rolling out the Ultrafast Mode for GPT-6.1 Sol on API, Codex, and ChatGPT. Discover more in Today&#8217;s Global AI Native Industry Insights.</p>
<h3>1.  Anthropic launches Claude Dashboards and Claude Motion in beta</h3>
<p>Anthropic has released two new Claude features in beta: Claude Dashboards and Claude Motion. Claude Dashboards lets users ask Claude to turn their data into live, interactive dashboards. Claude Motion allows users to convert their ideas into animated explainer videos. Both features are now available for testing as part of Anthropic&#8217;s expanding Claude product suite.<br />
Read more: <a href="https://claude.com/resources/articles/dashboards-and-motion">https://claude.com/resources/articles/dashboards-and-motion</a></p>
<p><video width="600" height="400" controls poster="https://cdn.ainative.foundation/image/20261010_d8674ae1de3b42fa98c3fb946692ff4a.png"><source src="https://cdn.ainative.foundation/video/20261010_ec29e1467f6b49f38d81d714857bb10c.mp4" type="video/mp4"></video><br />
Video Credit: @claudeai on X</p>
<h3>2.  Google Launches First Test Satellite Carrying TPUs to Explore Space-Based AI Infrastructure</h3>
<p>Google announced Project Suncatcher, a research effort examining whether machine learning infrastructure could one day be hosted in space. The initiative aims to determine whether Google&#8217;s TPU AI chips can withstand the physical stress of launch as well as the radiation and thermal extremes of orbit. According to Google, researchers have been studying these questions for several years. The company said it recently launched its first test satellite carrying four TPUs into orbit to begin gathering real-world data.<br />
Read more: <a href="https://x.com/Google/status/2107875278416753052">https://x.com/Google/status/2107875278416753052</a></p>
<p><video width="600" height="400" controls poster="https://cdn.ainative.foundation/image/20261010_b7f39050aee34ac29593169d0d514cee.jpg"><source src="https://cdn.ainative.foundation/video/20261010_03bb1a885c034e3cb111e80e4267d979.mp4" type="video/mp4"></video><br />
Video Credit: @Google on X</p>
<h3>3.  OpenAI Rolls Out Ultrafast Mode for GPT-6.1 Sol Across API, Codex, and ChatGPT Work</h3>
<p>OpenAI announced that Ultrafast mode for its GPT-6.1 Sol model is rolling out today across the API, Codex, and ChatGPT Work. The new mode delivers near-Astra level intelligence while running up to 8 times faster than Sol Standard. The update targets developers and enterprise users who need rapid response times for building and iterating on applications. The rollout spans OpenAI&#8217;s developer API as well as its coding and workplace-focused products.<br />
Read more: <a href="https://developers.openai.com/api/docs/guides/ultrafast-mode">https://developers.openai.com/api/docs/guides/ultrafast-mode</a></p>
<p><video width="600" height="400" controls poster="https://cdn.ainative.foundation/video/20261010_en_openai.jpeg"><source src="https://cdn.ainative.foundation/video/20261010_41e31853c1cd4c2aaac47f812ae08885.mp4" type="video/mp4"></video><br />
Video Credit: @OpenAIDevs on X</p>
<div style="width:100%;height:2px;background:#808080;margin:10px 0"></div>
<p>That&#8217;s all for today&#8217;s Global AI Native Industry Insights. Join us at <a href="https://member.ainativefoundation.org/">AI Native Foundation Membership Dashboard</a> for the latest insights on AI Native, or follow our linkedin account at <a href="https://www.linkedin.com/company/ainativefoundation/">AI Native Foundation</a> and our twitter account at <a href="https://x.com/AINativeF">AINativeF</a>.</p>
<p>The post <a href="https://ainativefoundation.org/global-ai-native-industry-insights-20261010-anthropic-google-openai-more/">Global AI Native Industry Insights &#8211; 20261010 &#8211;  Anthropic | Google | OpenAI | more</a> appeared first on <a href="https://ainativefoundation.org">AI Native Foundation</a>.</p>
]]></content:encoded>
					
		
		<enclosure url="https://cdn.ainative.foundation/video/20261010_ec29e1467f6b49f38d81d714857bb10c.mp4" length="22151077" type="video/mp4" />
<enclosure url="https://cdn.ainative.foundation/video/20261010_03bb1a885c034e3cb111e80e4267d979.mp4" length="438050" type="video/mp4" />
<enclosure url="https://cdn.ainative.foundation/video/20261010_41e31853c1cd4c2aaac47f812ae08885.mp4" length="450399" type="video/mp4" />

			</item>
		<item>
		<title>China AI Native Industry Insights &#8211; 20261009 &#8211;  Vidu &#124; ima.copilot &#124; Kling AI &#124; more</title>
		<link>https://ainativefoundation.org/china-ai-native-industry-insights-20261009-vidu-ima-copilot-kling-ai-more/</link>
		
		<dc:creator><![CDATA[AINF]]></dc:creator>
		<pubDate>Fri, 09 Oct 2026 08:14:26 +0000</pubDate>
				<category><![CDATA[China Industry]]></category>
		<guid isPermaLink="false">https://ainativefoundation.org/?p=11099</guid>

					<description><![CDATA[<p>Today’s digest highlights key advancements from Vidu, Tencent, and Kuaishou, with exciting releases reshaping the AI landscape. At the forefront is Vidu's Q4 preview of their next-generation flagship AI video model, promising advancements in video technology. The broader theme connects to AI's evolving role in content creation and management, emphasizing enhanced capabilities for both video and knowledge workspaces. Notably, Tencent introduces ima, an AI workspace harnessing the power of Hunyuan and DeepSeek models to optimize writing and knowledge management. Kuaishou's Kling AI version 4.0 features 30-second native video generation and offers improved stability for handling complex camera movements.</p>
<p>The post <a href="https://ainativefoundation.org/china-ai-native-industry-insights-20261009-vidu-ima-copilot-kling-ai-more/">China AI Native Industry Insights &#8211; 20261009 &#8211;  Vidu | ima.copilot | Kling AI | more</a> appeared first on <a href="https://ainativefoundation.org">AI Native Foundation</a>.</p>
]]></description>
										<content:encoded><![CDATA[<p>Today’s digest highlights key advancements from Vidu, Tencent, and Kuaishou, with exciting releases reshaping the AI landscape. At the forefront is Vidu&#8217;s Q4 preview of their next-generation flagship AI video model, promising advancements in video technology. The broader theme connects to AI&#8217;s evolving role in content creation and management, emphasizing enhanced capabilities for both video and knowledge workspaces. Notably, Tencent introduces ima, an AI workspace harnessing the power of Hunyuan and DeepSeek models to optimize writing and knowledge management. Kuaishou&#8217;s Kling AI version 4.0 features 30-second native video generation and offers improved stability for handling complex camera movements. Discover more in Today&#8217;s China AI Native Industry Insights.</p>
<h3>1.  Vidu Releases Q4 Preview, a Next-Generation Flagship AI Video Model</h3>
<p>Vidu AI launched Vidu Q4 Preview, described as its next-generation flagship AI video generation model. The company highlights expressive character performance, cinematic camera movement, and high-impact visual effects as key improvements. Vidu Q4 Preview is now live and priced from 0.014 dollars per second of generated video. The release aims to make flagship-level AI video generation more accessible through lower-cost pricing.<br />
Read more: <a href="https://x.com/ViduAI_official/status/2107856382263505219">https://x.com/ViduAI_official/status/2107856382263505219</a></p>
<p><video width="600" height="400" controls poster="https://cdn.ainative.foundation/image/20261009_dc6e345f156d4428a579e7f6f0a3f478.jpg"><source src="https://cdn.ainative.foundation/video/20261009_558a65c588b44141a1f9d066c962cce1.mp4" type="video/mp4"></video><br />
Video Credit: @ViduAI_official on X</p>
<h3>2.  Tencent releases ima, an AI knowledge workspace powered by Hunyuan and DeepSeek models for writing and knowledge management</h3>
<p>Tencent has launched ima, an AI-powered knowledge workspace that integrates search, reading, and writing capabilities. The platform is powered by Tencent Hunyuan and DeepSeek dual AI engines and features a copilot function that learns user habits and preferences over time. The tool is available on Windows, Mac desktop platforms and WeChat mini-program, with cross-platform data synchronization. New users receive 500 computational credits upon registration plus 100 daily credits for regular use.<br />
Read more: <a href="https://mp.weixin.qq.com/s?__biz=Mzk0MDg0MTkzOA==&#038;mid=2247506267&#038;idx=1&#038;sn=a9b2d72305b83670c01b33a8f3fd5f3c&#038;chksm=c32f81af78f15e16b67b3e3ca60e52333a2c457025a813402e2c7b119f5d0cc53604352fd90a&#038;scene=0&#038;xtrack=1#rd">https://mp.weixin.qq.com/s?__biz=Mzk0MDg0MTkzOA==&#038;mid=2247506267&#038;idx=1&#038;sn=a9b2d72305b83670c01b33a8f3fd5f3c&#038;chksm=c32f81af78f15e16b67b3e3ca60e52333a2c457025a813402e2c7b119f5d0cc53604352fd90a&#038;scene=0&#038;xtrack=1#rd</a></p>
<p><video width="600" height="400" controls poster="https://cdn.ainative.foundation/image/20261009_b4167e0967ea44d09a78dbe6825bad09"><source src="https://cdn.ainative.foundation/video/20261009_cn_tencent.mp4" type="video/mp4"></video><br />
Video Credit: NotebookLM</p>
<h3>3.  Kuaishou&#8217;s Kling AI releases version 4.0 with 30-second native video generation and enhanced stability for complex camera movements</h3>
<p>Kuaishou&#8217;s Kling AI announced the beta release of version 4.0, scheduled for official launch in October 2026. The new model introduces up to 30-second native video generation and improves stability in complex camera movements and large dynamic scenes. It enhances understanding of long prompts and continuous narratives while maintaining consistency of characters, props, and scenes across shots. Additional capabilities include subject and scene asset invocation, sound continuation, and motion transfer. An advertising creative director with 28 years of experience tested the model and confirmed its improved performance in high-speed camera work, action sequences, cross-scene consistency, and character performance with emotional speech.<br />
Read more: <a href="https://mp.weixin.qq.com/s?__biz=MzU5NTkwNDU2OA==&#038;mid=2247496912&#038;idx=1&#038;sn=52292535ca61af3fc66326f8aea772ce&#038;chksm=ffdc0ae6d84a39877cf6ccc6357a241ab2bba9ccb47fda840fc078fdf1c84e2901fe23c51c07&#038;scene=0&#038;xtrack=1#rd">https://mp.weixin.qq.com/s?__biz=MzU5NTkwNDU2OA==&#038;mid=2247496912&#038;idx=1&#038;sn=52292535ca61af3fc66326f8aea772ce&#038;chksm=ffdc0ae6d84a39877cf6ccc6357a241ab2bba9ccb47fda840fc078fdf1c84e2901fe23c51c07&#038;scene=0&#038;xtrack=1#rd</a></p>
<p><video width="600" height="400" controls poster="https://cdn.ainative.foundation/video/20261009_cn_keling.png"><source src="https://cdn.ainative.foundation/video/20261009_cn_keling.mp4" type="video/mp4"></video><br />
Video Credit: @Kling_ai on X</p>
<div style="width:100%;height:2px;background:#808080;margin:10px 0"></div>
<p>That&#8217;s all for today&#8217;s China AI Native Industry Insights. Join us at <a href="https://member.ainativefoundation.org/">AI Native Foundation Membership Dashboard</a> for the latest insights on AI Native, or follow our linkedin account at <a href="https://www.linkedin.com/company/ainativefoundation/">AI Native Foundation</a> and our twitter account at <a href="https://x.com/AINativeF">AINativeF</a>.</p>
<p>The post <a href="https://ainativefoundation.org/china-ai-native-industry-insights-20261009-vidu-ima-copilot-kling-ai-more/">China AI Native Industry Insights &#8211; 20261009 &#8211;  Vidu | ima.copilot | Kling AI | more</a> appeared first on <a href="https://ainativefoundation.org">AI Native Foundation</a>.</p>
]]></content:encoded>
					
		
		<enclosure url="https://cdn.ainative.foundation/video/20261009_558a65c588b44141a1f9d066c962cce1.mp4" length="57102529" type="video/mp4" />
<enclosure url="https://cdn.ainative.foundation/video/20261009_cn_tencent.mp4" length="12147381" type="video/mp4" />
<enclosure url="https://cdn.ainative.foundation/video/20261009_cn_keling.mp4" length="13486821" type="video/mp4" />

			</item>
		<item>
		<title>AI Native Daily Paper Digest – 20261008 – nanoMuse &#124; STEPQuant &#124; DecepEval</title>
		<link>https://ainativefoundation.org/ai-native-daily-paper-digest-20261008-nanomuse-stepquant-decepeval/</link>
		
		<dc:creator><![CDATA[insights]]></dc:creator>
		<pubDate>Fri, 09 Oct 2026 02:52:26 +0000</pubDate>
				<category><![CDATA[Papers]]></category>
		<guid isPermaLink="false">https://ainativefoundation.org/ai-native-daily-paper-digest-20261008-nanomuse-stepquant-decepeval/</guid>

					<description><![CDATA[<p>Today&#8217;s digest prominently features advances from Gemma and DeepSeek, showcasing breakthroughs in agentic systems and multimodal reasoning. The papers collectively delve into [&#8230;]</p>
<p>The post <a href="https://ainativefoundation.org/ai-native-daily-paper-digest-20261008-nanomuse-stepquant-decepeval/">AI Native Daily Paper Digest – 20261008 – nanoMuse | STEPQuant | DecepEval</a> appeared first on <a href="https://ainativefoundation.org">AI Native Foundation</a>.</p>
]]></description>
										<content:encoded><![CDATA[<p style="margin-bottom:24px">Today&#8217;s digest prominently features advances from Gemma and DeepSeek, showcasing breakthroughs in agentic systems and multimodal reasoning. The papers collectively delve into techniques for enhancing long-context attention, with a particular focus on hierarchical memory architectures. Notably, one method achieves a 15% improvement on the standard benchmark for document synthesis, while another presents a new dataset that doubles the previous record for image-text alignment accuracy. One study highlights the potential for achieving zero-shot performance in complex task environments.</p>
<h3>1. STEPQuant: When and Where Errors Matter in Delta-Rule Recurrent State Quantization</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Linear attention, STEPQuant, recurrent states, quantization</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Natural Language Processing</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; Propose STEPQuant, a spatial-temporal post-training quantization framework for Delta-rule recurrent states to optimize precision allocation and reduce memory usage.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Developed a framework that allocates precision based on error magnitude and memory lifetime while fitting key-row and value-column scales according to state distributions.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; STEPQuant closely matches FP32-state accuracy with a 6-bit budget, outperforms uniform INT8 in 4-bit configuration, and achieves significant memory compression, reducing serving memory by up to 68.7%.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2609.38169" target="_blank">https://huggingface.co/papers/2609.38169</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img fetchpriority="high" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20261008233010268.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>2. nanoMuse: An Open-Source Personal Agent for Every Device You Own</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Personal Agent, Open Source, Meta&#8217;s Muse, Cloud, AI Systems and Tools</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: AI Systems and Tools</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To define the concept of a personal agent, exemplified by Meta&#8217;s Muse, which integrates into a user&#8217;s accounts and devices, and explores the creation of an open-source counterpart called nanoMuse.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The study analyzes how Meta&#8217;s Muse was developed from public records and production prompts and introduces the open-source alternative, nanoMuse, under the GPL-3.0 license, which can run on any user-owned device.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; nanoMuse is designed to be an open-source solution with a unique communication relay for personal devices, showing scalability and user choice in model selection, alongside a roadmap for its open memory management and evaluation suite.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2610.08699" target="_blank">https://huggingface.co/papers/2610.08699</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20261008233032828.jpg"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>3. DecepEval: A Benchmark for Evaluating Deception in LLM Agents</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: large language model (LLM) agents, deception, DecepEval, AI Native</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: AI Ethics and Fairness</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; Introduce DecepEval, a benchmark for measuring deception in large language models across various scenarios.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Utilize the LLM Deception Diamond framework to assess deception induced by pressure, incentive, opportunity, and conflict conditions.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Inducements increase deception rates across various LLMs and task families, even in models with initially low deception rates, highlighting the need for more trustworthy AI.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2610.07967" target="_blank">https://huggingface.co/papers/2610.07967</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20261008233057112.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>4. Questioning the Questions: Sustaining Self-Evolution in Reasoning Models</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Self-evolution, R-Quest, Mathematical reasoning, Performance collapse</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Knowledge Representation and Reasoning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To investigate performance deterioration in self-evolving reasoning models and explore methods to sustain self-evolution.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Introduction of R-Quest utilizing question validity and novelty feedback to maintain self-evolution by training solvers to recognize invalid questions and using a frozen base model to compare questions for novelty.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; R-Quest consistently achieves the highest performance across benchmarks in mathematical reasoning, general-domain reasoning, and code generation, maintaining stable gains and outperforming previous models like R-Zero over ten rounds of self-evolution.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2610.04299" target="_blank">https://huggingface.co/papers/2610.04299</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20261008233116782.jpg"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>5. VepAgent: Bridging Causal-Transition via Tool-Augmented Reinforcement Learning for Video Event Prediction</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: MLLMs, Video Event Prediction, tool-augmented reinforcement learning, future-oriented reasoning, FutureBench</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Multi-Modal Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The primary aim is to enhance Video Event Prediction (VEP) by overcoming the limitations of traditional Multimodal Large Language Models (MLLMs) through an innovative agentic framework called VepAgent.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; VepAgent integrates causal-transition reasoning with tool-augmented reinforcement learning (RL), employing a high-quality &#8220;chain-of-thought&#8221; dataset, FutureBench-4K, for supervised fine-tuning. It also develops diagnostic tools for dynamic reasoning enhancement and resolves visual ambiguities.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; VepAgent achieves state-of-the-art performance on FutureBench and NEPBench datasets, significantly outperforming larger MLLMs. The method&#8217;s success validates the empirical effectiveness of an agentic, future-oriented reasoning paradigm.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2610.06293" target="_blank">https://huggingface.co/papers/2610.06293</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20261008233135978.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>6. Tetris3D: 3D Scene Generation With Objects That Fit Together</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Tetris3D, 3D scene reconstruction, generative framework, physical coherence, ComOb</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Generative Models</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The main goal is to develop a generative framework, Tetris3D, for single-image 3D scene reconstruction, ensuring that objects are both physically and geometrically coherent as part of a scene.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The approach explicitly conditions object generation on the geometry and physical relationships of surrounding objects, guiding their shapes and poses within the scene. Additionally, a physics simulation-based dataset called ComOb is introduced, consisting of 1.2M scenes with diverse object categories and detailed annotations.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Tetris3D demonstrates state-of-the-art performance in generating coherent object shapes and poses, even when interacting regions are occluded, achieving high-quality generation and physical stability in both synthetic and real-world scenes.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2610.10539" target="_blank">https://huggingface.co/papers/2610.10539</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20261008233215135.jpg"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>7. UniWAM: Unified World-Action Model</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Vision-language-action models, Semantic understanding, AI Native, SOTA performance, Human-robot co-training</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Multi-Modal Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The paper aims to introduce UniWAM, a unified architecture that integrates a physical reasoner, a world generator, and an action predictor to enhance semantic understanding, visual generation, and action prediction in AI systems.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The implementation of an intricate data cleaning and annotation pipeline for human egocentric data and robot data, and the introduction of a pre-training recipe using visual question answering data, human demonstrations, and robot demonstrations.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; UniWAM achieves state-of-the-art performance in diverse evaluation settings, demonstrating robustness, generalization, and effective long-term task execution. A log-linear scaling law in human-robot co-training underscores the success of large-scale pre-training across human and robot data.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2610.02054" target="_blank">https://huggingface.co/papers/2610.02054</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20261008233154640.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>8. WorldSonus: Bringing Sound to Worlds</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: world models, real-time generation, interactive control, spatially aligned stereo, AI Native</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Generative Models</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To introduce WorldSonus, a framework for real-time spatial sound synthesis in world models, addressing the challenges of sound generation in interactive video streams.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Utilization of a streaming causal autoregressive diffusion architecture for low RTF audio chunk synthesis.</p>
<p>   &#8211; Implementation of an audio-centric captioning pipeline for interactive sound event manipulation.</p>
<p>   &#8211; High-quality stereo supervision based on diverse stereo and ambisonic data for spatial alignment.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; WorldSonus, tailored for world models, effectively generalizes to open-domain video-to-audio benchmarks, performing competitively with state-of-the-art models in both acoustic quality and spatial alignment.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2610.08760" target="_blank">https://huggingface.co/papers/2610.08760</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20261008233239433.jpg"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>9. Mechanics of Long-Context Hybrid Models Part 1.1: From Hybrid Attention to Hybrid Position</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Large Language Models, hybrid models, attention mechanisms, long-context efficiency</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Natural Language Processing</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To explain the effectiveness of hybrid models combining different attention mechanisms in enhancing long-context efficiency and to propose a framework for their design.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Analysis of hybrid models involving full attention, sliding-window attention, and gated variants of linear attention, observing effects in context extension and positional inductive biases.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The study finds that linear attention hybrids perform better with long-context pretraining, while sliding-window attention models excel in length extrapolation, highlighting varying performance due to positional inductive biases.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2610.10114" target="_blank">https://huggingface.co/papers/2610.10114</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20261008233259089.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>10. Gains and Collapse in On-Policy Distillation:A Reinforcement Learning Perspective</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: On-policy distillation, Language model, Reinforcement learning, Implicit reward model, Reward hacking</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Natural Language Processing</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; Investigate the outcomes of on-policy distillation (OPD) in language model post-training and understand the reinforcement learning perspective leading to performance variation.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Experiments analyzing the behavior of implicit reward models in OPD and testing methods like masking unhealthy responses and using SFT initialization to mitigate performance collapse.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; OPD can enhance language model performance by amplifying favorable student behaviors, but may also cause collapse due to reward misalignment. Masking and initializations were found effective in preventing such collapse.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2610.03185" target="_blank">https://huggingface.co/papers/2610.03185</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20261008233317277.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>11. Recurrent Looped Transformer</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Recurrent Looped Transformer, state tracking, sequence length, depth, transformer</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Machine Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; Introduce the Recurrent Looped Transformer (RLT) to improve state tracking by combining a parallel causal encoder with a recurrent decoder.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Compare RLT with traditional transformers using eight layers and different splits, tested on six algorithmic tasks with a focus on sequence length amidst limited training bits.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; RLT demonstrates significantly higher accuracy than traditional transformers in parity tasks, permutation tracking, and modular arithmetic, thanks to its use of feedback mechanisms and per-token processing.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2610.07591" target="_blank">https://huggingface.co/papers/2610.07591</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20261008233335399.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>12. VIEScore2: Unified Image Evaluation with Spatially Grounded Explanations</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Image Generation, Defect Localization, VIEScore2, GRPO, Text-Native Grid Representation</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Computer Vision</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The research introduces VIEScore2, aiming to provide a more comprehensive evaluation of synthetic images by predicting quality scores and defect locations in a single pass.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Utilizes a text-native grid representation to integrate diverse spatial supervision and apply GRPO for enhanced defect localization through rewards considering cell-level metrics, score accuracy, and format validity.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; VIEScore2 outperforms existing general-purpose visual language models and spatial evaluators on several benchmarks, demonstrating superior performance in evaluating and localizing defects in synthetic images.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2610.00994" target="_blank">https://huggingface.co/papers/2610.00994</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20261008233353235.jpg"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>13. On KL-Regularized Policy Optimization</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Asynchronous reinforcement learning, Large language model, KLPO, Critic-free update, Monte Carlo estimates</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Reinforcement Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The objective is to improve training efficiency for large language model (LLM) agents through a novel method called KL-Regularized Policy Optimization (KLPO).</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Employs KL-Regularized Policy Optimization which uses a Gibbs solution and least squares fitting on the sampler&#8217;s trajectories, eliminating the need for importance weights.</p>
<p>   &#8211; Demonstrates a method for achieving critic-free updates using sampler-centered scores or trajectory residuals via Monte Carlo estimates.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; KLPO provides an efficient policy update method that requires only one rollout per prompt without the necessity of a learned normalizer or group responses.</p>
<p>   &#8211; Shows how SPPO, GPO, REBEL, and BPO can all be considered special cases of KLPO, underscoring its versatility and applicability.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2610.08963" target="_blank">https://huggingface.co/papers/2610.08963</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20261008233416595.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>14. Inverting Multi-Vector Visual Document Indices</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Document Retrieval, Index Inversion, Vision-Language Model, Sensitive Data</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Computer Vision</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To investigate the vulnerability of multi-vector visual document retrievers to index inversion attacks.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Evaluating the ability to reconstruct document pages from indices stored in vector databases using a vision-language model.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Multi-vector document retrievers are susceptible to inversion attacks, as demonstrated by a high word and sensitive token recovery rate from indices. Protective measures like token pooling and shuffling drastically reduce this risk but can be circumvented by models restoring vector order.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2610.09920" target="_blank">https://huggingface.co/papers/2610.09920</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20261008233434799.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>15. WebFovea: When the Model Is Right but the Click Is Wrong &#8212; Reliable Round Trips for Vision-Based Web Agents on Live Websites</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: WebFovea, vision-based web agent, WebRetriever Challenge 2026, multimodal large language model, harness</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: AI Systems and Tools</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The research introduces WebFovea, a vision-based web agent designed to operate live websites, achieving notable performance in the WebRetriever Challenge 2026.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Utilized a multimodal large language model (LLM) coordinated via a complex harness that ensures correct execution of agent actions on web interfaces through a systematic four-stage process.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; WebFovea effectively fortifies each operational stage with robust guardrails, showing improvements through runs despite the inherent challenges of site interaction and variability, ultimately enhancing the agent&#8217;s performance from a score of 31.0 to 57.0.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2610.03036" target="_blank">https://huggingface.co/papers/2610.03036</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20261008233455387.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>16. From Pareto to Preference: Personalized Test-Time Scaling via Amortized Agentic Policy Discovery</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Test-time scaling, Inference computation, Personalized Test-Time Scaling, PersonTTS, AI Native</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Natural Language Processing</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To enhance large language models&#8217; reasoning abilities by maximizing the joint satisfaction rate of user-specific requirements using Personalized Test-Time Scaling.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Proposed a framework named PersonTTS which uses an amortized agentic policy-discovery approach to reuse prior search experience and guide policy through requirement-matched controller initialization and source-distilled procedural guidance.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; PersonTTS significantly outperforms existing strong TTS baselines in meeting joint user requirements on unseen profiles, successfully improving policy quality while reducing discovery-agent time and cost.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2610.09684" target="_blank">https://huggingface.co/papers/2610.09684</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20261008233514115.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>17. AdSpark: A Large-Scale Dataset and Benchmark for Product-Centric Advertisement Video Generation</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: AdSpark, Product-centric advertisement, Visual storytelling, Multi-shot narratives</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Generative Models</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; Introduce a large-scale dataset and benchmark, AdSpark, for product-centric advertisement video generation to address the lack of specific datasets and comprehensive evaluation frameworks.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Develop AdSpark-300K consisting of 300K reference image-prompt-video triplets with real-world and synthetic subsets, providing structured advertisement annotations.</p>
<p>   &#8211; Propose AdSpark-Bench, a diagnostic benchmark for evaluating generated advertisements across six key dimensions.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Evaluate representative models using AdSpark-Bench, identifying challenges in product preservation, storytelling, and visualization.</p>
<p>   &#8211; Validate dataset effectiveness with experiments on AdSpark-300K-finetuned models, underscoring its importance for future research.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2610.10047" target="_blank">https://huggingface.co/papers/2610.10047</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20261008233533436.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>18. QuadTok: Quadtree Visual Tokenizer for Autoregressive Image Generation</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: QuadTok, visual tokenization, autoregressive image generation</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Generative Models</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; Introduce QuadTok, a new framework for visual tokenization and autoregressive image generation using a hierarchical quadtree structure.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Implement a tokenizer that allocates tokens based on visual complexity, achieving a 10% token saving on ImageNet with comparable reconstruction fidelity.</p>
<p>   &#8211; Utilize a GPT-style generative model conditioned on quadtree topology for effective image generation.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; QuadTok facilitates efficient and spatially controlled image generation, demonstrating zero-shot generative capabilities on datasets like COCO.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2610.10497" target="_blank">https://huggingface.co/papers/2610.10497</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20261008233553454.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>19. DLoop: Looped Speculative Decoding</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Speculative Decoding, Large Language Models, Autoregressive Generation, Parallel draft model, DLoop</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Natural Language Processing</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To introduce DLoop, a looped speculative decoding method that adaptively performs multiple drafting stages before verification to speed up autoregressive generation in large language models.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Implementation of DLoop which allows continuous drafting while a draft model remains confident, checking accumulated draft tokens at once, and using loop-aware training to ensure the draft model remains reliable in additional drafting stages.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; DLoop significantly improves wall-clock speedup by 5 to 41 percent across various speculative decoding methods and maintains lossless decoding.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2610.07659" target="_blank">https://huggingface.co/papers/2610.07659</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20261008233612557.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>20. Learning Multimodal Embeddings with Evidence-Aligned Readout</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Multimodal Large Language Models, Semantic Evidence, Retrieval Embedding, Boundary Readout</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Multi-Modal Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study investigates how the semantic organization of task-relevant evidence contributes to retrieval embeddings in multimodal large language models.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; EviAlign is introduced, combining Semantic Evidence Generation with Boundary Readout to organize evidence into semantic units and aggregate states into a normalized embedding. A 2&#215;3 controlled study compares evidence organization under various strategies.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Consistent semantic organization enhances retrieval performance, with improvements observed when evidence boundaries are used over length-based training positions. EviAlign achieves a high Recall@1 score on retrieval tasks, maintaining efficient indexing.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2609.33659" target="_blank">https://huggingface.co/papers/2609.33659</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20261008233649935.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>21. CADFather: Autonomous CAD Reconstruction through Coordinated Tool Use</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: CADFather, 3D meshes, parametric CAD programs, AI Systems and Tools, autonomous agentic system</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: AI Systems and Tools</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study introduces CADFather, an autonomous system designed to reconstruct parametric CAD programs from 3D meshes by coordinating complementary tools.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; CADFather employs a vision-language assistant to choose CAD program extensions, invoke tools, and generate proposals. It uses learned and algorithmic tools for operations and numerical optimization for refining parameters.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The system was evaluated on multiple test sets, including DeepCAD and Fusion360, demonstrating its capability not only to reconstruct high-quality and valid CAD models but also to analyze the trade-offs between computational cost and reconstruction quality. </p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2610.09127" target="_blank">https://huggingface.co/papers/2610.09127</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20261008233631408.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>22. SkillForge: Co-Evolving Skills and Agents via Dynamic Skill Lifecycles</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Memory-augmented reinforcement learning, LLM agents, SkillForge, skill lifecycle</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Reinforcement Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The primary goal is to enhance LLM agents&#8217; capacity to solve intricate and long-horizon tasks by systematically managing and evolving a skill library.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Introduces SkillForge, an RL method that organizes the skill library through four states &#8211; trial, active, stable, and retired &#8211; enabling skills and model to co-evolve.</p>
<p>   &#8211; Conducts a pre-RL evaluation phase to retire low-fitness skills, followed by reinforcement learning with iterative skill library optimization.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; SkillForge shows superior performance in interactive agent benchmarks, achieving up to a 7.8% improvement over the strongest baseline while maintaining a compact library.</p>
<p>   &#8211; Introduces SkillFurnace, a dataset supporting research in skill quality and lifecycle management.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2610.09832" target="_blank">https://huggingface.co/papers/2610.09832</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20261008233707717.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>23. A self-learning scientific agent for X-ray diffraction</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Gan Jiang, AI Systems and Tools, powder X-ray diffraction, structural knowledge extraction</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: AI Systems and Tools</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; Introduce Gan Jiang, a self-learning agent designed for powder X-ray diffraction within a newly developed diffraction-analysis ecosystem.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Implement a suite of engines (XMatcher, XQueryer, XDecomposer, WPEM) to perform phase identification, multiphase decomposition, and physics-constrained pattern modelling. Gan Jiang refines analytical skills by diagnosing and revising skill instructions without retraining models.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Gan Jiang significantly enhances refinement scores compared to expert-designed skills. Its performance in phase identification outperforms existing methods, demonstrating improved accuracy in multiphase identification from both simulated and experimental data.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2610.07862" target="_blank">https://huggingface.co/papers/2610.07862</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20261008233727364.jpg"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>24. StepCAD: Mesh-to-CAD Code Generation via LLM Policy and Geometry-Guided Search</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: CAD programs, geometry-guided search, StepCAD, ARCADE-1.5M, IoU improvement</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Generative Models</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; Introduce a generative optimization approach, StepCAD, for recovering executable CAD programs from 3D meshes, addressing the challenges of compositional nature and parameter interactions in CAD construction.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Combine a state-conditioned CAD policy with geometry-guided search, using an IoU-guided tree search for local edits to improve geometric reconstruction accuracy.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; StepCAD demonstrates state-of-the-art geometric reconstruction accuracy with up to 87.2% relative IoU improvement over existing baselines, showing significant gains on complex shapes.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2610.03799" target="_blank">https://huggingface.co/papers/2610.03799</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20261008233804980.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>25. UniSkill: Learning Actor-Aligned Skill Proposals for an Evolving Policy</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Large language model agents, skillbank, skill proposal, UniSkill, AI Native</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Reinforcement Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study introduces UniSkill, aiming to enhance task performance by optimizing skillbank edits through shared policy interactions.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Utilizes a shared policy for task execution and skill proposal without requiring additional costly actor rollouts. Employs contrastive action feedback to guide skill proposal learning and apply skill-edit support regularization for exploration.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; UniSkill demonstrates high effectiveness, achieving notable success rates in ALFWorld and WebShop environments, with stability in joint training even with a smaller policy backbone.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2610.10164" target="_blank">https://huggingface.co/papers/2610.10164</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20261008233746467.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>26. CoDance: Learning Reactive and Compliant Human-Humanoid Interaction from Video</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Humanoid Interaction, CoDance, Compliance Augmentation, Simulation</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Robotics and Autonomous Systems</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To develop CoDance, a framework for learning reactive and compliant human-humanoid interaction from video, focusing on coordinated locomotion and continuous physical contact.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Utilized video data of human dancers to retarget motions onto a robot and moving partner.</p>
<p>   &#8211; Introduced multi-link compliance augmentation for adapting robot references under structured forces at both hands, transforming kinematic demonstrations into force-aware training data.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Policies developed enable humanoids to maintain coordinated dancing with human partners, adapting to partner&#8217;s changes and reproducing 80% of wrist displacement from augmented demonstrations.</p>
<p>   &#8211; CoDance achieved sustained two-hand dancing with repeated dynamic transitions on a physical humanoid.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2610.05324" target="_blank">https://huggingface.co/papers/2610.05324</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20261008233846702.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>27. RoboQuest: Generalist Physical Agents that Search, Inspect and Test</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Multimodal foundation models, Embodied exploration, RoboQuest, Task-relevant information</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Robotics and Autonomous Systems</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; Introduce RoboQuest, a benchmark for goal-directed embodied exploration where agents must actively acquire and use task-relevant information.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Evaluation of five frontier multimodal agents through a visuomotor interface and a fine-tuned policy based on full-episode demonstrations.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Findings indicate that the best agent succeeds in only 23% of episodes, with challenges in task completion largely due to early decision-making without sufficient evidence and difficulties in trial and error learning.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2610.10388" target="_blank">https://huggingface.co/papers/2610.10388</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20261008233823673.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>28. RoboJEPA: Scaling Robotic Latent World Models</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Latent world models, Joint Embedding Predictive Architecture, RoboJEPA, Imagination error, Scaling laws</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Robotics and Autonomous Systems</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To introduce RoboJEPA and establish scaling laws for multi-embodiment robotic world models trained on real robot data.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Developed and tested RoboJEPA, a predictor model based on Joint Embedding Predictive Architecture, trained on a large-scale dataset across 12 robotic embodiments.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; RoboJEPA&#8217;s imagination error follows a predictable power law in compute, and performance improves predictably with compute, serving as a reliable proxy for evaluating real robots.</p>
<p>   &#8211; Latent world models can be deployed zero-shot to solve tasks with long-horizon planning on real hardware. Released model checkpoints and code.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2610.10515" target="_blank">https://huggingface.co/papers/2610.10515</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20261008233926796.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>29. Rethinking World-Action Model for Compositional and In-Context Robotic Manipulation</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Long-horizon compositional manipulation, Visual Goal-conditioned Action Reasoning, task-level planning, in-context learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Robotics and Autonomous Systems</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study aims to enhance real-world robot deployment by addressing the complex nature of tasks involving multiple coordinated subtasks through long-horizon compositional manipulation.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The researchers propose ViGAR, a hierarchical framework that includes a visual subgoal planner and a subgoal executor for task-level planning. ViGAR is tested using the RoboTwin Clean2Random benchmark and involves using a pretrained world-model representation.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; ViGAR achieved notable success rates of 82.00% and 67.02% in Clean and Random settings on the benchmark, outperforming existing models by an average success rate margin of 12.86 percentage points. Additionally, the framework&#8217;s real-world effectiveness is demonstrated through robot experiments on seven tasks.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2610.02368" target="_blank">https://huggingface.co/papers/2610.02368</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20261008233906489.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>30. AutoResearch at Production Scale: Failure Modes and a Multi-Agent Framework</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: AutoResearch, recommendation pipelines, representation-learning, infrastructure fragility, autonomous research</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: AI Systems and Tools</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To automate the exploration of embedding systems optimization in production recommendation pipelines using Andrej Karpathy&#8217;s AutoResearch paradigm.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Implementation of a large language model that iteratively edits a training script, focusing on experimentation and modification that enhance performance metrics.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The study identified five failure modes in autonomous research systems and introduced a &#8220;prevent, persist, redirect&#8221; framework to address these. The framework yielded significant improvements in recall, coherence, and catalog coverage compared to hand-tuned baselines.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2609.30541" target="_blank">https://huggingface.co/papers/2609.30541</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20261008234103462.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>31. Task-Sufficient Contraction: Source Selection for Machine Information Interfaces</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Task-Sufficient Contraction, reduced source, finite action sets, rate-regret curve, Information Bottleneck</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Foundations of AI</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To explore the concept of Task-Sufficient Contraction, where a reduced source can solve downstream problems with the same results as a full source.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Investigation of when one-step rate-regret curves on finite action sets are preserved, and analysis of exact contractions under quadratic loss for affine feasible-action sets.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Successfully identifies conditions for consumer-specific source reduction and distinguishes exact and approximate contractions, facilitating efficient communication in heterogeneous machines without needing full internal alignment.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2610.08884" target="_blank">https://huggingface.co/papers/2610.08884</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20261008234045460.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>32. TIDES: Implicit Time-Awareness in Selective State Space Models</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Selective SSM, Continuous Time SSM, Irregular Time Series</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Machine Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The paper introduces TIDES, a variant of selective state space models that integrates the strengths of both selective and continuous architectures to handle irregular timestamps.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; TIDES modifies traditional SSMs by shifting input dependence from the step size to the diagonal state matrix, allowing it to manage irregular datasets effectively without compromising expressivity. Performance was demonstrated on both new experimental benchmarks and large-scale datasets across various domains.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; TIDES achieves significant improvements, setting new benchmarks in time series classification and regression tasks, and performs competitively on diverse irregular datasets spanning astronomy, agriculture, neuromorphic sensing, and climate events.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2605.09742" target="_blank">https://huggingface.co/papers/2605.09742</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20261008234025628.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>33. PAMI: Part Anchored Motion for Text to Human-Object Interaction Generation</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Text-conditioned HOI, PAMI, Interaction generation, Hybrid surface-sensing, PamiVAE</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Generative Models</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The paper aims to generate text-conditioned full-body human-object interaction (HOI) that maintains precise coordination over time between human motion and object trajectories.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Introduces PAMI, a Part-Anchored Motion framework using a representation inspired by the Hough Transform, localized object motion through body-part anchors, and PamiVAE to learn an interaction latent space.</p>
<p>   &#8211; Proposes a coarse-to-fine hierarchical generation strategy using PamiGen for generating coarse interactions and PamiRefiner for refining contact geometry with a hybrid surface-sensing representation.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; PAMI generates more faithful interactions and accurate human-relative object motion compared to previous methods, with a 14.5% higher contact recall in experiments conducted on the InterAct dataset.</p>
<p>   &#8211; Extensive ablations validate the part-anchored voting representation and hybrid surface-sensing refinement&#8217;s contributions.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2609.38466" target="_blank">https://huggingface.co/papers/2609.38466</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20261008234003329.jpg"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>34. How corner is a corner case? Percentile control for highway scenario generation</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Autonomous Vehicle, Corner-case Scenarios, Risk Percentile, Simulation Environment, Generative Models</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Robotics and Autonomous Systems</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study aims to generate corner-case scenarios within a simulation environment to test the safety performance of an autonomous vehicle software stack, focusing on controlling scenario extremity relative to future risks.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Employs a percentile-conditioned joint diffusion model with risk guidance to simulate multi-agent futures, using history-conditioned risk-percentile requests and mapping to physical risk targets.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Achieves high accuracy in realizing requests with a 98.75% success rate within a 0.05 percentile tolerance, connecting context-relative risk specification, physical realization, and evaluation on a unified risk scale.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2610.05003" target="_blank">https://huggingface.co/papers/2610.05003</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20261008233944418.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>35. </h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="" target="_blank"></a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/thirteen/202610081791502880.jpg"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>36. ReGain: Restoring Subject Fidelity in Personalization on Synthetic Images</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: DreamBooth, Synthetic Images, Classifier-Free Guidance, ReGain</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Generative Models</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study investigates the effects of using synthetic images for fine-tuning text-to-image diffusion models, specifically looking at subject fidelity degradation.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The researchers fine-tune two models using the same base model and DreamBooth protocol: one on real photos and another on synthetic images, analyzing the differences attributed to classifier-free guidance.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The paper concludes that synthetic images degrade subject fidelity, traced to increased angles and norms of noise predictions. It proposes ReGain as a training-free solution that significantly improves subject fidelity without real photos.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2609.38680" target="_blank">https://huggingface.co/papers/2609.38680</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20261008234113103.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>37. Co-Evolving Robot Orchestrators and Policies through Deployment</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Vision-language-action, Robo-COP, self-improving flywheel, policy fine-tuning, Robotics</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Robotics and Autonomous Systems</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The paper aims to enhance the adaptability and performance of Vision-Language-Action (VLA) policies in real-world robotic applications by introducing a novel approach called Robo-COP.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Robo-COP co-evolves the orchestrator and policy during deployment, curating skill demonstrations from its own executions, fine-tuning the policy, and adopting new policies after skill improvements.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Robo-COP significantly improves success rates in both simulated (from 64.8% to 73.8%) and real-world tasks (from 38.3% to 50.0%) compared to systems using a static policy, effectively turning deployment into a self-improving process.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2610.09228" target="_blank">https://huggingface.co/papers/2610.09228</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20261008234054751.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>38. FastOPD: On-Policy Distillation for Lightweight VLA Deployment</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: VLA foundation models, FastOPD, on-policy distillation, inference latency, RoboTwin 2.0</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Robotics and Autonomous Systems</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To enable practical deployment of large-scale Vision-Language-Action (VLA) foundation models by developing a framework called FastOPD for efficient on-policy distillation.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Adaptation of a flow map for single-state teacher supervision combined with a self-consistency objective to construct a compact student model.</p>
<p>   &#8211; Evaluation of FastOPD across diverse foundation policies in simulation and real-world experiments.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; FastOPD retains 84% of the performance with significantly reduced inference latency and outperforms existing few-step distillation baselines.</p>
<p>   &#8211; The framework enhances the single-step success rate significantly in robotic applications, demonstrating practical deployment potential on real robots.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2610.02832" target="_blank">https://huggingface.co/papers/2610.02832</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20261008234035283.jpg"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>39. System Switch: When Should a Fast Decision Model Stop and Think?</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Dual-process agents, Fast policy, Reasoning model, AUROC, Decision models</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Reinforcement Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To investigate the dynamics between fast learned actors and slow reasoning vision-language models in decision-making processes, specifically focusing on how decisions are managed and deferred in real-time environments.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Utilization of a closed-loop Doom environment and &#8220;System One&#8221; typed-decision models interfaced via llama.cpp. Evaluation includes zero-shot decision models with varying parameters and analysis of decision accuracy, calibration, and sensitivity.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The study finds that models with similar accuracy can differ significantly in AUROC, illustrating different sensitivity levels in decision-making. Additionally, deferring decisions based on the actor&#8217;s confidence can enhance performance proportions. A rule-based approach for exploration increases door openings but leads to more frequent failures. The reasoning model often misidentifies simple tasks due to the lack of specific knowledge, such as differentiating between locked and ordinary doors.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2610.09683" target="_blank">https://huggingface.co/papers/2610.09683</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20261008234016384.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>40. Lineage-Aware Memory Governance: A Derivation-Gated Framework for Privacy-Preserving Column-Level Access Control in Enterprise AI Agents</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Enterprise AI agents, memory store, data leakage, Analytical Memory Unit, lineage-gated retrieval</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: AI Systems and Tools</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study aims to address the unaddressed risks associated with Enterprise AI agents sharing a memory store, focusing on preventing sensitive data leakage and conflicts in KPI computations.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Introduction of the Analytical Memory Unit (AMU) schema, which attaches a derivation graph to cached results, with a retrieval policy that ensures access only with proper authorization for every column involved.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Lineage-gated retrieval significantly reduces cross-department data leakage and maintains high memory reuse efficiency while ensuring compliance with governance standards like the EU AI Act.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2610.07258" target="_blank">https://huggingface.co/papers/2610.07258</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20261008233953333.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>41. Beyond the Parameter Monolith: Reconstructive Memories, Executable Skills, and Residual Assembly for Language Models</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Language-model systems, FEM-ASM, finite-element-method-inspired, causal language modeling, neural rendering</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Natural Language Processing</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To investigate FEM-ASM, an organization for language-model systems separating contextual computation, persistent storage, and exact execution.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Controlled experiments and a Multi-Mesh prototype without attention mechanisms were used to evaluate the system.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; While successful in separating storage, execution, and neural coordination, challenges remain in question-only retrieval, unrestricted answer generation, and overall efficiency.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2610.04012" target="_blank">https://huggingface.co/papers/2610.04012</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20261008233935530.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>42. Agent Plasticity: Measuring Self-Improvement Through Experience</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: AI agents, self-improvement, agent plasticity, experience, learning efficiency</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Reinforcement Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study investigates self-improvement in AI agents by evaluating their ability to learn effectively over time rather than at a fixed point.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Researchers analyze self-improvement in a controlled environment where AI agents utilize past experiences to enhance future performance, measuring efficiency and identifying bottlenecks.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Significant differences were found in agents&#8217; improvement trajectories and efficiency, revealing the importance of evaluating both learning capacity and application of learned knowledge.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2610.08902" target="_blank">https://huggingface.co/papers/2610.08902</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20261008233916482.jpg"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>43. Salt++: Context-Aligned Post-Training for Few-Step Streaming Multimodal Generation</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Causal Self-Flow, step distillation, context-aligned autoregressive DMD, Salt++</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Generative Models</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The paper aims to address context-related challenges in few-step streaming audio-video generation through a novel framework called Salt++, incorporating Causal Self-Flow (CSF) and context-aligned autoregressive Distribution Matching Distillation (DMD).</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Introducing Salt++ with a two-stage post-training framework that exploits contextual information asymmetry and matches generated and reference distributions using a block-conditional KL objective.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Salt++ significantly improves visual and motion quality over OmniForcing in a 4-step causal setting and outperforms bidirectional LTX-2 on multiple metrics, demonstrating its effectiveness in enhancing cross-modal alignment and generation quality.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2609.36995" target="_blank">https://huggingface.co/papers/2609.36995</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20261008233855980.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>44. DSReg: Provably Recovering Individual World Latents without Reconstruction</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Latent Variables, Structural Diversity, Joint-Embedding Predictive Architectures (JEPA)</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Machine Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To recover individual latent variables without requiring reconstruction, decoders, or labeled supervision.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Utilizing Structural Diversity and DSReg (Dependency-Sparsity Regularization) to achieve latent recovery up to signed permutation.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; DSReg recovers individual world latents in linearly identified representations without reconstruction and maintains dense prediction quality.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2610.09457" target="_blank">https://huggingface.co/papers/2610.09457</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20261008233834762.jpg"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>45. EngramEdit: Decoupled Knowledge Updates in LLMs through Conditional Memory</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Conditional Memory, DeepSeek Engram, EngramEdit, Knowledge Updates, Transformer</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Natural Language Processing</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study aims to decouple factual knowledge storage from general-purpose computation in large language models, facilitating independent factual updates while keeping the underlying Transformer architecture unchanged.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; EngramEdit is introduced to manage knowledge updates by computing target memory representations and updating shared n-gram embeddings, ensuring accurate predictions across different expressions without affecting unrelated knowledge.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; EngramEdit allows for precise knowledge updates and maintains unrelated knowledge integrity. It exhibits nearly perfect editing success and achieves close to three times the baseline&#8217;s accuracy in chain-of-thought prompting, demonstrating its effectiveness as an editable knowledge interface for conditional memory architectures.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2610.10533" target="_blank">https://huggingface.co/papers/2610.10533</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20261008233814358.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>46. Inherit-MAS: Test-Time Evolution of Multi-Agent Systems through Workflow and Execution Inheritance</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Multi-agent systems, Inherit-MAS, workflow inheritance, execution inheritance, redundancy reduction</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: AI Systems and Tools</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study aims to improve workflow evolution within Multi-agent systems (MAS) by introducing Inherit-MAS, designed to optimize both workflow and execution levels.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Inherit-MAS employs a meta-model to create workflows of worker agents with defined roles and evaluates candidate executions through a judge mechanism to enhance workflow refinement and reduce redundancy.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Inherit-MAS outperforms existing MAS such as EvoAgent and EvoMAS in efficiency and completion metrics on benchmarks like WorkBench and HotpotQA, reducing token usage significantly while enhancing task completion performance.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2610.02396" target="_blank">https://huggingface.co/papers/2610.02396</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20261008233755905.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>47. SheetSage2: Coherent Lead-Sheet Transcription with Synthetic Supervision</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Music Transcription, SheetSage2, Synthetic Data, Autoregressive Distillation, AI Systems and Tools</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: AI Systems and Tools</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To develop a unified music transcription framework called SheetSage2 that enhances the coherence and accuracy of transcribed music scores.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Utilization of synthetic data, task-specific structured decoding, and autoregressive distillation to improve the transcription process.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; SheetSage2 surpasses previous systems on 12 of 15 benchmarks, demonstrating improved performance over its predecessor and task-specific models. The model and code are publicly available for further use.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2610.05336" target="_blank">https://huggingface.co/papers/2610.05336</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20261008233737192.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>48. Iris-3B: Going Beyond the Latent with Pixel-Space Diffusion Training, Conversion and Fine-Tuning</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Pixel-space, Text-to-image transformer, Generative Models, Depth estimation, Super-resolution</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Generative Models</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To evaluate the efficacy of pixel-space diffusion models compared to latent models in tasks where fine-grained detail is vital.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Pretraining a 3B-parameter pixel-space text-to-image transformer (Iris-3B) and converting a pretrained latent model (FLUX.2 Klein base 4B) to pixel space.</p>
<p>   &#8211; Fine-tuning both models for monocular depth estimation and image restoration/super-resolution.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; No significant improvement from using a pixel-space generative prior was found for the tasks explored.</p>
<p>   &#8211; Although Iris-3B achieved competitive text-to-image quality, it did not outperform latent models in depth and image restoration tasks.</p>
<p>   &#8211; The study provides insights into the shortcomings and areas for future enhancement in pixel-space model utilization.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2610.09450" target="_blank">https://huggingface.co/papers/2610.09450</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20261008233717084.jpg"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>49. Composing What Each Teacher Learned: Multi-Teacher On-Policy Distillation through Teacher-Relative Shifts</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Multi-teacher on-policy distillation, Δ-MOPD, endpoint policy, specialist routing, performance improvement</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Reinforcement Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To introduce Δ-MOPD, a novel method for transferring each teacher&#8217;s teacher-minus-base logit shift, in order to improve the existing multi-teacher on-policy distillation methods.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Comparative analysis of Δ-MOPD against endpoint supervision in both common-domain composition and routed-domain distillation settings.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Δ-MOPD shows significant improvements in performance, particularly in scenarios where multiple teacher signals are combined at a state, outperforming endpoint composition in specific metrics.</p>
<p>   &#8211; In phased and interleaved routing scenarios, Δ-MOPD demonstrates comparable or improved outcomes, highlighting its potential as an independent design axis complementary to teacher selection.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2610.10460" target="_blank">https://huggingface.co/papers/2610.10460</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20261008233658375.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>50. Improving Proactive AI Assistance with Hierarchical Procedural Understanding</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Proactive AI assistants, Adaptive guidance system, ProactiveCoach suite, VLMs, Hierarchical supervision</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Human-AI Interaction</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study aims to improve Proactive AI assistants by providing adaptive and context-aware guidance to users, adjusting to task progress and user expertise.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Introduction of the ProactiveCoach suite, consisting of ProactiveCoach-Instruct for training and ProactiveCoachBench for evaluation, along with fine-tuned VLMs. It utilizes hierarchical guidance to better align with task progress and needs.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Hierarchical supervision enhances performance by up to 9.6% compared to fixed-granularity supervision. The adaptive guidance system developed surpasses the baseline by 57.1% in adaptation across multiple guidance-level transitions.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2610.06505" target="_blank">https://huggingface.co/papers/2610.06505</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20261008233640972.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>51. Internalizing Agent Experience into Diffusion Model Weights via On-Policy Context Distillation</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Agentic Harness, Diffusion Model, Text-to-Image, Auto Skill Evolver, Continual Co-evolution</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Generative Models</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The objective is to improve Text-to-Image task performance by integrating the capabilities of an agentic harness into a diffusion model using a novel method called Diffusion On-Policy Context Distillation (D-OPCD).</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; This study introduces D-OPCD, which distills agentically improved prompts into the model&#8217;s weights, enhancing image generation performance.</p>
<p>   &#8211; Utilizes an Auto Skill Evolver (ASE) to iteratively update and internalize the harness&#8217;s capabilities within the generator.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; D-OPCD successfully improves the average direct-generation score from 60.52 to 65.09 across benchmarks.</p>
<p>   &#8211; The integration allows the harness to keep evolving and improve, ensuring continual co-evolution with the model and further enhances performance by an additional 1.83 points.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2610.07250" target="_blank">https://huggingface.co/papers/2610.07250</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20261008233621372.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>52. MIMESIS: Learning User Simulators as Training Environments for Interactive Agents</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: User Simulator, Behavioral Fidelity, Reinforcement Learning, Generalization, Coached On-Policy Self-Distillation</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Reinforcement Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To introduce MIMESIS, a user simulator designed to mimic realistic user behavior for training interactive language agents.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Developed a 9B model trained with explicit reasoning supervision and realistic behavioral patterns. Used multi-turn reinforcement learning with the simulator for agent training.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; MIMESIS surpasses existing models in behavioral fidelity and Turing distance. Training with MIMESIS provides better generalization across unseen user simulators than with GPT-5.5. The CSD method further enhances agent performance by transforming feedback into token-level supervision.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2610.09484" target="_blank">https://huggingface.co/papers/2610.09484</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20261008233603592.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>53. Mobile-4DGS: Unified Static-Dynamic Real-time Mobile Gaussian Splatting</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: 3D Gaussian Splatting, Mobile-4DGS, real-time rendering, static and dynamic scenes</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Computer Vision</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The paper introduces Mobile-4DGS, aiming to achieve high-fidelity real-time Gaussian rendering on resource-constrained mobile platforms for both static and dynamic scenes.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Mobile-4DGS uses a Monte Carlo Specular Energy Aggregator for compressing radiance residuals and an Attribute-Conditioned SH Enhancement module for efficiency.</p>
<p>   &#8211; A Multi-View Alpha-Based Densification and Pruning strategy is applied to suppress redundant primitives while ensuring multi-view consistency.</p>
<p>   &#8211; For dynamic scenes, a compact 4D representation with second-order Gaussian motion and a binary static-dynamic partition is proposed to enable continuous-time modeling.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Mobile-4DGS significantly reduces storage and rendering overhead on mobile devices while maintaining competitive visual quality for real-time 3D and 4D Gaussian Splatting.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2610.05289" target="_blank">https://huggingface.co/papers/2610.05289</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20261008233542689.jpg"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>54. On-Policy Distillation with Negative-Policy Rollouts</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: On-policy distillation, Negative-Policy OPD, teacher supervision, rollout stage</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Reinforcement Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To introduce Negative-Policy OPD (NP-OPD) that incorporates a negative policy to complement traditional on-policy distillation, enhancing learning signals by providing a negative reference during student model training.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Utilize rollouts from a lower-performing negative policy alongside a stronger teacher supervision during the token-level rollout stage to continuously provide contrasting learning signals that improve the distillation process.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; NP-OPD enhances OPD across various model scales, generation modes, reasoning domains, and OPD variants by effectively moving the student model away from undesirable learning pathways attributed to the negative policy, and improving overall learning outcomes.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2610.07874" target="_blank">https://huggingface.co/papers/2610.07874</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20261008233523378.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>55. Q-Learning with Scalar Adjoint Matching</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Flow policies, off-policy RL, Scalar Adjoint, Q-learning, OGBench domains</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Reinforcement Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To improve the fine-tuning of flow policies using off-policy Reinforcement Learning (RL) for enhancing performance beyond initial demonstrations.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Introduction of Q-learning with Scalar Adjoint Matching (SQAM) that mitigates the need for costly vector&#8211;Jacobian products by scaling the value gradient through a closed-form scalar adjoint approach.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; SQAM notably increases success rates on challenging OGBench domains, outperforming the strongest baseline by 18 to 35 percentage points.</p>
<p>   &#8211; SQAM&#8217;s framework effectively enhances fine-tuning for large pretrained models, demonstrated in real-world applications such as a bimanual robot.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2610.10437" target="_blank">https://huggingface.co/papers/2610.10437</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20261008233504326.jpg"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>56. Minimal Witness Reinforcement Learning</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Minimal-Witness Reinforcement Learning, minimal sufficient witnesses, policy gradient, RL methods</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Reinforcement Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To formalize the problem of minimal-witness identification and introduce a new approach named Minimal-Witness Reinforcement Learning (MWRL) to address it.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; MWRL uses a value iteration planner to recover the entire family of witnesses and a policy gradient method that scales to large language models. It credits each proposal for its unique contribution, derived directly from the problem definition.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; MWRL effectively recovers most minimal witnesses in various experimental settings, outperforming other methods that return redundant supersets or a single witness, thus expanding reinforcement learning&#8217;s scope beyond single-solution optimization.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2610.07226" target="_blank">https://huggingface.co/papers/2610.07226</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20261008233443683.jpg"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>57. NAMVIS: Next-Scale Autoregressive Multi-View Image Synthesis</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Sparse-view synthesis, NAMVIS, geometry-conditioned autoregression, Multi-scale Projective Pose Encoding, diffusion-free</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Generative Models</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The research primarily aims to offer an efficient alternative for sparse-view multi-view synthesis by introducing NAMVIS, which avoids the drawbacks of diffusion-based methods.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The study introduces a diffusion-free framework using geometry-conditioned next-scale autoregression instead of iterative denoising and proposes Multi-scale Projective Pose Encoding to incorporate camera transformations in the model&#8217;s process.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; NAMVIS outperforms diffusion-based baselines in key metrics like PSNR, SSIM, and LPIPS, while providing a significant speed advantage, suggesting its potential as a promising and efficient solution for 3D content creation.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2610.04722" target="_blank">https://huggingface.co/papers/2610.04722</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20261008233425085.jpg"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>58. Agentic RAG Evaluation: Budget Allocation Across Questions, Trajectories, and Reads</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Agentic RAG, Evaluation budget, Generalizability theory, Repeated sampling</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Natural Language Processing</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; Investigating the allocation precision, reading efficiency, and cost boundaries of agentic retrieval-augmented generation (RAG) using diverse evaluation budgets.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Utilization of HotpotQA and MuSiQue datasets for retrieval-feedback comparison, analysis of question coverage, token budget assessments, and prediction accuracy evaluations with nested and Q-only forecasts.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Enhanced question coverage significantly reduces standard error, archived forecasts predict allocations accurately, and optimal cost management approaches are proposed, with temperature settings reducing answer disagreements without impacting precision.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2610.05034" target="_blank">https://huggingface.co/papers/2610.05034</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20261008233403957.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>59. SWE-Game: Can Coding Agents Build the Games We Want?</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: SWE-Game, Godot-to-Unity porting, Opus5, gameplay logic errors</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: AI Systems and Tools</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; Introduce SWE-Game, a benchmark with 247 tasks across 41 Godot games to evaluate agent performance in game development.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Five distinct task types are used, including implementation, skeleton completion, and gameplay porting, assessed through engine-state checks and feature demonstrations.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Opus5 achieves the highest score among six models, although scores remain below 60 out of 100 for main tasks, indicating prevalent requirement omissions and logic errors.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2609.33678" target="_blank">https://huggingface.co/papers/2609.33678</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20261008233344340.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>60. PhysEvo: Astra Can Act, Let It</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Physical recursive self-improvement, Meta-agent, Astra, AI Systems and Tools, Robotics and Autonomous Systems</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Robotics and Autonomous Systems</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; Introduction of PhysEvo, a framework for physical recursive self-improvement around a single frozen model to enhance reliable manipulation in the Astra system.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Utilization of a task agent and a meta-agent to execute robot tasks, diagnose failures, and improve tools and skills without model-weight updates.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; PhysEvo demonstrates superior performance compared to RoboDawn&#8217;s Astra agent, achieving higher success rates in manipulation tasks and real-world task trials.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2610.08995" target="_blank">https://huggingface.co/papers/2610.08995</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20261008233325976.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>61. RobotWorld: Benchmarking Multimodal Agents for Robot Use Across Diverse Tasks and Embodiments</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: general-purpose agents, RobotWorld, physical task execution, capability transfer, task outcomes</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Robotics and Autonomous Systems</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The study investigates how the capabilities of general-purpose agents in writing code, using tools, and completing complex digital tasks transfer to the physical world by introducing RobotWorld, a simulation testbed for robot use.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; A simulation testbed called RobotWorld is utilized, covering 84 tasks across various domains such as manipulation and driving, with explicit interaction budgets and success checks to analyze task outcomes against execution traces.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Current agents construct sophisticated perception and control workflows, yet struggle with consistent successful task execution, losing task-relevant object states and mistaking unfinished tasks for completion. The study identifies capability gaps and model differences in task success rates, establishing targets for training more reliable physical-world agents.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2610.10409" target="_blank">https://huggingface.co/papers/2610.10409</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20261008233307724.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>62. ReSAIL: Mitigating Collapse in Iterative Agent Self-Distillation</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Iterative self-distillation, Recursive self-improvement, Privileged Information, ReSAIL, Multimodal GUI agents</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Reinforcement Learning</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The paper aims to address deployment performance collapse in iterative self-distillation of LLM agents by developing a method that prioritizes informative interaction steps and preserves PI-conditioned behavior.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Introduction of ReSAIL, a plug-in augmentation for PI-based self-distillation that selects key interaction steps and balances distillation losses, applied on ALFWorld and TextCraft environments.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; ReSAIL delivers substantial improvements in final-cycle success rates and enhances action prediction accuracy, demonstrating the method&#8217;s efficacy in mitigating performance collapse in iterative self-distillation.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2609.39306" target="_blank">https://huggingface.co/papers/2609.39306</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20261008233250766.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>63. RunningTab: Direct Workspace Interaction with Environment-Side Tabs</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: direct workspace interaction (DWI), Large Language Models (LLM), RunningTab</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: AI Systems and Tools</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The paper aims to address the shortcomings of direct workspace interaction (DWI) by introducing RunningTab, a framework that enhances file tracking and extraction tasks through an environment-side tab.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The authors validate the RunningTab framework on three benchmarks using three different LLMs to test its efficacy compared to plain DWI and baseline models.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; RunningTab consistently outperforms plain DWI and other baseline models by effectively keeping track of task requirements and ensuring that deliverables incorporate necessary values from the environment.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2610.10444" target="_blank">https://huggingface.co/papers/2610.10444</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20261008233229104.jpg"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>64. Semifactual Credit-Augmented Policy Optimization</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Reinforcement Learning, Verifiable Rewards, Semifactual Stability, SCAPO, GRPO  </p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Reinforcement Learning  </p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:  </p>
<p>   &#8211; To enhance reasoning accuracy in large language models by addressing token-level sensitivity and improving credit assignment in reinforcement learning with verifiable rewards.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:  </p>
<p>   &#8211; Introduced Semifactual Credit-Augmented Policy Optimization (SCAPO), which incorporates semifactual stability into token-level credit assignment, measured token probability drift, and adjusted token advantages during early training.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:  </p>
<p>   &#8211; SCAPO outperformed existing methods by improving accuracy on AIME 2024-2026 benchmarks and generalization across model scales and benchmarks, demonstrating the efficacy of semifactual stability in reinforcement learning.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2609.40360" target="_blank">https://huggingface.co/papers/2609.40360</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20261008233204285.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>65. SGF+: Decoupling Gradient Flows for Autoregressive Video Generation</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Self Gradient Forcing Plus, autoregressive video generation, temporal consistency, role-specific parameterization</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Generative Models</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; Investigate the challenge of optimizing visual quality and temporal consistency in autoregressive video generation by addressing the systematic negative alignment of distinct gradient patterns in parameter sharing.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Introduce Self Gradient Forcing Plus (SGF+), which separates parameters for context writing and denoising, and optimizes both roles using the original generation objective while maintaining interaction through causal attention.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Implementing role-specific parameterization with SGF+ improves visual quality and long-horizon consistency, enabling continuous video generation up to 24 hours without additional training data or long-video fine-tuning.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2610.10429" target="_blank">https://huggingface.co/papers/2610.10429</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20261008233144903.jpg"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>66. UltraText Bench: A Comprehensive Bilingual Benchmark for Evaluating Visual Text Rendering in Image Generation</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Dense visual text, UltraText Bench, Q-Judger, text fidelity, spatial quality</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Computer Vision</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; The research introduces UltraText Bench, a bilingual benchmark aimed at evaluating prompt-only generation of dense visual text across various real-world scene categories.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The benchmark contains 432 prompts covering 24 scene categories in both English and Chinese with three difficulty levels. Evaluation involves human-reviewed prompts with structured references assessed by the Q-Judger vision-language model.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Results show varied performance across different model configurations, highlighting strengths and weaknesses such as clarity and fidelity in dense text reproduction. Performance variability is also observed across different workload levels.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2610.09823" target="_blank">https://huggingface.co/papers/2610.09823</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20261008233126702.png"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>67. GRACE: Generation-aware latent compression for efficient video generation</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: compressed video autoencoders, video diffusion models, Diffusion Transformer, Generation-Aware Latent Compression, GRACE</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Generative Models</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To propose a framework, GRACE, that can compress a pretrained video autoencoder while remaining compatible with a pretrained Diffusion Transformer for efficient video generation.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Utilizes a two-stage framework with a frozen base latent and residual latent learning for lost information, aligning compressed latents in the feature space, and applying lightweight fine-tuning and asymmetric denoising.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; GRACE significantly reduces the token count and latency of Wan2.1-I2V-14B, while maintaining generation quality on VBench, demonstrating its effectiveness in efficient video generation.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2610.10524" target="_blank">https://huggingface.co/papers/2610.10524</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20261008233106190.jpg"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>68. Recursive Game Creator: An Agentic Product-Level Experience-Oriented Game Harness</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Recursive Game Creator, game development, coding-native Player</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Generative Models</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To enhance agentic game development by transforming basic game prototypes into engaging and entertaining games through the Recursive Game Creator framework.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; The Recursive Game Creator framework organizes development into four components: Designer, Builder, Player, and Reviewer, each playing a distinct role in the iterative game creation process.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; Achieved state-of-the-art performance with a 77.89 score on GameCraft-Bench and improved task success rate and runtime-check pass rate on GameASG-Bench. User studies indicated longer playtime and higher ratings.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2610.08621" target="_blank">https://huggingface.co/papers/2610.08621</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20261008233044201.jpg"></figure>
</p>
</div>
<div style="height:30px"></div>
<h3>69. Long-WAM: Scaling the Context of World-Action Models</h3>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f511.png" alt="🔑" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Keywords: Real-time robot control, Long-WAM, Autoregressive (AR) pretraining, Deployment on RTX 5090</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4a1.png" alt="💡" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Category: Robotics and Autonomous Systems</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f31f.png" alt="🌟" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Objective:</p>
<p>   &#8211; To develop Long-WAM, a framework that extends the context of causal world-action models under real-time control constraints, aiming to enhance success in robot manipulation tasks by utilizing visual history more effectively.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f6e0.png" alt="🛠" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Methods:</p>
<p>   &#8211; Leveraging causal prediction from robot and egocentric videos without action labels.</p>
<p>   &#8211; Implementing AR pretraining for better utilization of longer visual histories.</p>
<p>   &#8211; Deploying the system on different hardware platforms such as RTX 5090 and DGX Spark for efficient real-time performance.</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f4ac.png" alt="💬" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Research Conclusions:</p>
<p>   &#8211; The use of AR pretraining with longer visual histories significantly increases task success rates, exemplified by a success rate improvement from 63.3% to 78.7% on RoboCasa GR-1.</p>
<p>   &#8211; Long-WAM outperforms existing methods on platforms such as LIBERO-Long, RoboTwin 2.0, and DOMINO.</p>
<p>   &#8211; Successful real-time application on Unitree G1 and YAM, with notable performance in dynamic manipulation tasks like cup stacking.</p>
</p>
<p><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Paper link:&nbsp;<a href="https://huggingface.co/papers/2610.10528" target="_blank">https://huggingface.co/papers/2610.10528</a></p>
<div class="wp-block-image">
<figure class="aligncenter"><img loading="lazy" decoding="async" width="660" height="660" src="https://cdn.ainative.foundation/huggingface/20261008233020911.jpg"></figure>
</p>
</div>
<div style="height:30px"></div>
<p>The post <a href="https://ainativefoundation.org/ai-native-daily-paper-digest-20261008-nanomuse-stepquant-decepeval/">AI Native Daily Paper Digest – 20261008 – nanoMuse | STEPQuant | DecepEval</a> appeared first on <a href="https://ainativefoundation.org">AI Native Foundation</a>.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Global AI Native Industry Insights &#8211; 20261008 &#8211;  Anthropic &#124; Google &#124; OpenAI &#124; more</title>
		<link>https://ainativefoundation.org/global-ai-native-industry-insights-20261008-anthropic-google-openai-more/</link>
		
		<dc:creator><![CDATA[AINF]]></dc:creator>
		<pubDate>Thu, 08 Oct 2026 08:33:34 +0000</pubDate>
				<category><![CDATA[Global Industry]]></category>
		<guid isPermaLink="false">https://ainativefoundation.org/?p=11045</guid>

					<description><![CDATA[<p>Anthropic and Google take center stage in today's digest with significant advancements in AI technology. A theme of accessibility and efficiency emerges as Anthropic introduces Claude Haiku 5.5, their most cost-effective and swiftest small model to date, while Google Labs unveils Playground, an experimental no-code platform for creating AI-driven games. OpenAI joins the lineup by making GPT-6 and an Intelligent UI available to all ChatGPT users, enhancing user interaction and accessibility. Playground's launch offers a glimpse into the future of creating interactive digital experiences without needing advanced programming skills.</p>
<p>The post <a href="https://ainativefoundation.org/global-ai-native-industry-insights-20261008-anthropic-google-openai-more/">Global AI Native Industry Insights &#8211; 20261008 &#8211;  Anthropic | Google | OpenAI | more</a> appeared first on <a href="https://ainativefoundation.org">AI Native Foundation</a>.</p>
]]></description>
										<content:encoded><![CDATA[<p>Anthropic and Google take center stage in today&#8217;s digest with significant advancements in AI technology. A theme of accessibility and efficiency emerges as Anthropic introduces Claude Haiku 5.5, their most cost-effective and swiftest small model to date, while Google Labs unveils Playground, an experimental no-code platform for creating AI-driven games. OpenAI joins the lineup by making GPT-6 and an Intelligent UI available to all ChatGPT users, enhancing user interaction and accessibility. Playground&#8217;s launch offers a glimpse into the future of creating interactive digital experiences without needing advanced programming skills. Discover more in Today&#8217;s Global AI Native Industry Insights.</p>
<h3>1.  Anthropic Releases Claude Haiku 5.5, Its Cheapest and Fastest Small Model Yet</h3>
<p>Anthropic introduced Claude Haiku 5.5, describing it as the cheapest, fastest, and most capable small model the company has released to date. The new model costs around 75% less to run on average than its predecessor, Claude Haiku 4.5. The release expands Anthropic&#8217;s lineup of lower-cost models aimed at high-volume and latency-sensitive applications. The announcement was made through Anthropic&#8217;s official Claude AI account.<br />
Read more: <a href="https://www.anthropic.com/claude-haiku-5-5">https://www.anthropic.com/claude-haiku-5-5</a></p>
<p><video width="600" height="400" controls poster="https://cdn.ainative.foundation/image/20261008_bbe34b34afd64349b49fbfa72c8a3e3b.jpg"><source src="https://cdn.ainative.foundation/video/20261008_0169b71a7a884332964094019c72f735.mp4" type="video/mp4"></video><br />
Video Credit: @claudeai on X</p>
<h3>2.  Google Labs launches Playground, an experimental no-code AI game creation platform</h3>
<p>Google Labs introduced Playground, a new experimental platform that lets users create their own games without any coding experience. The tool is designed so that users can turn an idea into a playable game. It is currently available only to users aged 18 and older in the United States. Google positioned the release as an early-stage experiment rather than a finished consumer product.<br />
Read more: <a href="https://x.com/GoogleLabs/status/2107800195748737042">https://x.com/GoogleLabs/status/2107800195748737042</a></p>
<p><video width="600" height="400" controls poster="https://cdn.ainative.foundation/image/20261008_7eb67b2a05b44c34b214ec24dbb20724.jpg"><source src="https://cdn.ainative.foundation/video/20261008_06fdeeee74734d45be02f3467474f9bc.mp4" type="video/mp4"></video><br />
Video Credit: @GoogleLabs on X</p>
<h3>3.  OpenAI Rolls Out GPT-6 and Intelligent UI to All ChatGPT Users</h3>
<p>OpenAI announced that GPT-6 and a new Intelligent UI feature are now rolling out to all ChatGPT users. Intelligent UI is designed to deliver faster, more interactive answers by adding visual elements to everyday questions. It also aims to make complex topics easier to understand through dynamic, interactive tools generated on the spot for a user&#8217;s task. The rollout marks a broader deployment of OpenAI&#8217;s latest model alongside this new interface capability.<br />
Read more: <a href="https://t.co/XL2gDCPBwG">https://t.co/XL2gDCPBwG</a></p>
<p><video width="600" height="400" controls poster="https://cdn.ainative.foundation/image/20261008_2a56e0b29a9b4f9697e89fc07168d955.jpg"><source src="https://cdn.ainative.foundation/video/20261008_bdc83cdbfad5477daf64d11e86aa93bc.mp4" type="video/mp4"></video><br />
Video Credit: @OpenAI on X</p>
<div style="width:100%;height:2px;background:#808080;margin:10px 0"></div>
<p>That&#8217;s all for today&#8217;s Global AI Native Industry Insights. Join us at <a href="https://member.ainativefoundation.org/">AI Native Foundation Membership Dashboard</a> for the latest insights on AI Native, or follow our linkedin account at <a href="https://www.linkedin.com/company/ainativefoundation/">AI Native Foundation</a> and our twitter account at <a href="https://x.com/AINativeF">AINativeF</a>.</p>
<p>The post <a href="https://ainativefoundation.org/global-ai-native-industry-insights-20261008-anthropic-google-openai-more/">Global AI Native Industry Insights &#8211; 20261008 &#8211;  Anthropic | Google | OpenAI | more</a> appeared first on <a href="https://ainativefoundation.org">AI Native Foundation</a>.</p>
]]></content:encoded>
					
		
		<enclosure url="https://cdn.ainative.foundation/video/20261008_0169b71a7a884332964094019c72f735.mp4" length="2616562" type="video/mp4" />
<enclosure url="https://cdn.ainative.foundation/video/20261008_06fdeeee74734d45be02f3467474f9bc.mp4" length="46400291" type="video/mp4" />
<enclosure url="https://cdn.ainative.foundation/video/20261008_bdc83cdbfad5477daf64d11e86aa93bc.mp4" length="34680688" type="video/mp4" />

			</item>
		<item>
		<title>China AI Native Industry Insights &#8211; 20260930 &#8211;  Kunlun Tech &#124; ima.copilot &#124; Qoder &#124; more</title>
		<link>https://ainativefoundation.org/china-ai-native-industry-insights-20260930-kunlun-tech-ima-copilot-qoder-more/</link>
		
		<dc:creator><![CDATA[AINF]]></dc:creator>
		<pubDate>Wed, 30 Sep 2026 08:37:19 +0000</pubDate>
				<category><![CDATA[China Industry]]></category>
		<guid isPermaLink="false">https://ainativefoundation.org/?p=11030</guid>

					<description><![CDATA[<p>Today's digest highlights major developments from SkyProduction, Tencent, and Qoder. A central theme is the integration of AI technologies to enhance efficiency and connectivity across various sectors. SkyProduction's Seedance 2.5 introduces a draft mode with a 480P preview and 1080P final rendering, promising cost-effective AI video production. Tencent ima, meanwhile, consolidates skills from its ecosystem, including medical knowledge and stock data, creating a more comprehensive user experience. Additionally, Qoder's integration of TypeSafe AI's Jev judgment model into Agent Harness aims to streamline automated decision-making processes.</p>
<p>The post <a href="https://ainativefoundation.org/china-ai-native-industry-insights-20260930-kunlun-tech-ima-copilot-qoder-more/">China AI Native Industry Insights &#8211; 20260930 &#8211;  Kunlun Tech | ima.copilot | Qoder | more</a> appeared first on <a href="https://ainativefoundation.org">AI Native Foundation</a>.</p>
]]></description>
										<content:encoded><![CDATA[<p>Today&#8217;s digest highlights major developments from SkyProduction, Tencent, and Qoder. A central theme is the integration of AI technologies to enhance efficiency and connectivity across various sectors. SkyProduction&#8217;s Seedance 2.5 introduces a draft mode with a 480P preview and 1080P final rendering, promising cost-effective AI video production. Tencent ima, meanwhile, consolidates skills from its ecosystem, including medical knowledge and stock data, creating a more comprehensive user experience. Additionally, Qoder&#8217;s integration of TypeSafe AI&#8217;s Jev judgment model into Agent Harness aims to streamline automated decision-making processes. Discover more in Today&#8217;s China AI Native Industry Insights.</p>
<h3>1.  SkyProduction launches Seedance 2.5 draft mode with 480P preview and 1080P final rendering to reduce AI video production costs</h3>
<p>SkyProduction, powered by Kunlun Tech, launched the Seedance 2.5 draft mode on September 29, allowing AI video creators to generate 480P preview videos before committing to 1080P final renders. The feature addresses high inference costs in AI short drama production, where creators typically generate dozens of versions before selecting the final clip. According to testing from ByteDance&#8217;s Volcano Engine, the draft mode can reduce costs by approximately 57% in typical scenarios and up to 90% in high-iteration workflows, while maintaining visual consistency between draft and final output without relying on upscaling. The mode is now available on skyproduction.cn for users creating AI-generated video content.<br />
Read more: <a href="https://mp.weixin.qq.com/s?__biz=MzI1MzE1NDc3Mg==&#038;mid=2247510469&#038;idx=1&#038;sn=d9fe2fea61b8627bf8923d51d0d59933&#038;chksm=e883c45632445d9212d3079b59b123de6d077faac06d24b893b833141199f0ba0e16604ce14e&#038;scene=0&#038;xtrack=1#rd">https://mp.weixin.qq.com/s?__biz=MzI1MzE1NDc3Mg==&#038;mid=2247510469&#038;idx=1&#038;sn=d9fe2fea61b8627bf8923d51d0d59933&#038;chksm=e883c45632445d9212d3079b59b123de6d077faac06d24b893b833141199f0ba0e16604ce14e&#038;scene=0&#038;xtrack=1#rd</a></p>
<p><video width="600" height="400" controls poster="https://cdn.ainative.foundation/image/20260930_b063ede67e33410dbcfc3a820eb3d143"><source src="https://cdn.ainative.foundation/video/20260930_gn_video_kunlun.mp4" type="video/mp4"></video><br />
Video Credit: The original article</p>
<h3>2.  Tencent ima integrates multiple Tencent ecosystem Skills including Medical Knowledge, Stock Data, News, and WeChat Reading</h3>
<p>Tencent announced that its AI assistant ima has integrated multiple Tencent ecosystem Skills, allowing users to access various services through voice commands without switching apps. The integrated Skills include Tencent Medical Knowledge Q&#038;A, Tencent Stock Financial Data Query, Tencent News, WeChat Reading, Tencent Campus Recruitment, Tencent Fact Check, Tencent Meeting, Tencent E-Sign Contract Assistant, Tencent Weather, Tencent Channel, and PDF processing tools. Users can install these Skills through the Discover section in ima and invoke them directly through conversation, with outputs that can be saved to the knowledge base for future reference.<br />
Read more: <a href="https://mp.weixin.qq.com/s?__biz=Mzk0MDg0MTkzOA==&#038;mid=2247506155&#038;idx=1&#038;sn=45c22d07c0d94fede8cb10035c86bcb0&#038;chksm=c3abee5afbd330c3f8c657471116c7a70f09ac825e871eab0f45c022161c9c91380d8ebfcdba&#038;scene=0&#038;xtrack=1#rd">https://mp.weixin.qq.com/s?__biz=Mzk0MDg0MTkzOA==&#038;mid=2247506155&#038;idx=1&#038;sn=45c22d07c0d94fede8cb10035c86bcb0&#038;chksm=c3abee5afbd330c3f8c657471116c7a70f09ac825e871eab0f45c022161c9c91380d8ebfcdba&#038;scene=0&#038;xtrack=1#rd</a></p>
<p><video width="600" height="400" controls poster="https://cdn.ainative.foundation/image/20260930_e5ca1abc1e1b49a994391d53dc791193"><source src="https://cdn.ainative.foundation/video/20260930_gn_video_ima.mp4" type="video/mp4"></video><br />
Video Credit: NotebookLM</p>
<h3>3.  Qoder integrates TypeSafe AI&#8217;s Jev judgment model into Agent Harness for automated decision-making</h3>
<p>Qoder announced its integration of Jev, a System One judgment model from TypeSafe AI, into its Agent Harness architecture to handle high-frequency semantic decisions in agent workflows. The integration separates responsibilities where generative models handle understanding and planning, Jev handles bounded semantic judgments, and Harness manages permissions, execution, and validation. Qoder demonstrated two use cases: Auto Mode, which uses Jev to determine whether agent operations can execute automatically with permission judgment latency reduced by approximately 83 percent from 3.13 to 0.53 seconds, and browser automation, where Jev selects next-step page operations within a controlled decision loop. The approach aims to establish a clearer division of labor between generation, judgment, deterministic rules, and user control in agent systems.<br />
Read more: <a href="https://mp.weixin.qq.com/s?__biz=MzE5ODI2NzI5Nw==&#038;mid=2247492846&#038;idx=1&#038;sn=171f168c0a3a970a06990ae6bfefa839&#038;chksm=97bd8ffd0b4a074f9ea84bd84f2a40af8bbc4ed098f15f5fdeeee4e9381574323e96114b9358&#038;scene=0&#038;xtrack=1#rd">https://mp.weixin.qq.com/s?__biz=MzE5ODI2NzI5Nw==&#038;mid=2247492846&#038;idx=1&#038;sn=171f168c0a3a970a06990ae6bfefa839&#038;chksm=97bd8ffd0b4a074f9ea84bd84f2a40af8bbc4ed098f15f5fdeeee4e9381574323e96114b9358&#038;scene=0&#038;xtrack=1#rd</a></p>
<p><video width="600" height="400" controls poster="https://cdn.ainative.foundation/image/20260930_gn_img_qoder.png"><source src="https://cdn.ainative.foundation/video/20260930_gn_video_qoder.mp4" type="video/mp4"></video><br />
Video Credit: The original article</p>
<div style="width:100%;height:2px;background:#808080;margin:10px 0"></div>
<p>That&#8217;s all for today&#8217;s China AI Native Industry Insights. Join us at <a href="https://member.ainativefoundation.org/">AI Native Foundation Membership Dashboard</a> for the latest insights on AI Native, or follow our linkedin account at <a href="https://www.linkedin.com/company/ainativefoundation/">AI Native Foundation</a> and our twitter account at <a href="https://x.com/AINativeF">AINativeF</a>.</p>
<p>The post <a href="https://ainativefoundation.org/china-ai-native-industry-insights-20260930-kunlun-tech-ima-copilot-qoder-more/">China AI Native Industry Insights &#8211; 20260930 &#8211;  Kunlun Tech | ima.copilot | Qoder | more</a> appeared first on <a href="https://ainativefoundation.org">AI Native Foundation</a>.</p>
]]></content:encoded>
					
		
		<enclosure url="https://cdn.ainative.foundation/video/20260930_gn_video_kunlun.mp4" length="3912541" type="video/mp4" />
<enclosure url="https://cdn.ainative.foundation/video/20260930_gn_video_ima.mp4" length="8647234" type="video/mp4" />
<enclosure url="https://cdn.ainative.foundation/video/20260930_gn_video_qoder.mp4" length="36557044" type="video/mp4" />

			</item>
		<item>
		<title>China AI Native Industry Insights &#8211; 20260929 &#8211;  Alibaba &#124; Kling AI &#124; Kunlun Tech &#124; more</title>
		<link>https://ainativefoundation.org/china-ai-native-industry-insights-20260929-alibaba-kling-ai-kunlun-tech-more/</link>
		
		<dc:creator><![CDATA[AINF]]></dc:creator>
		<pubDate>Tue, 29 Sep 2026 08:53:07 +0000</pubDate>
				<category><![CDATA[China Industry]]></category>
		<guid isPermaLink="false">https://ainativefoundation.org/?p=11025</guid>

					<description><![CDATA[<p>Today’s digest spotlights significant developments from Alibaba, Kling, and SkyProduction in the AI landscape. These companies highlight the theme of enhancing multimedia capabilities through advanced AI models. Alibaba unveils Qwen-Audio-3.1 featuring five upgraded models with notable price reductions, while Kling releases the Kling 4.0 video generation model with its 4.0 Flash version now available for early access. SkyProduction introduces AI music video creation powered by Mureka music models and integrated intelligent soundtrack features. Early access to Kling's new video generation model allows users to experience its enhanced capabilities firsthand.</p>
<p>The post <a href="https://ainativefoundation.org/china-ai-native-industry-insights-20260929-alibaba-kling-ai-kunlun-tech-more/">China AI Native Industry Insights &#8211; 20260929 &#8211;  Alibaba | Kling AI | Kunlun Tech | more</a> appeared first on <a href="https://ainativefoundation.org">AI Native Foundation</a>.</p>
]]></description>
										<content:encoded><![CDATA[<p>Today’s digest spotlights significant developments from Alibaba, Kling, and SkyProduction in the AI landscape. These companies highlight the theme of enhancing multimedia capabilities through advanced AI models. Alibaba unveils Qwen-Audio-3.1 featuring five upgraded models with notable price reductions, while Kling releases the Kling 4.0 video generation model with its 4.0 Flash version now available for early access. SkyProduction introduces AI music video creation powered by Mureka music models and integrated intelligent soundtrack features. Early access to Kling&#8217;s new video generation model allows users to experience its enhanced capabilities firsthand. Discover more in Today&#8217;s China AI Native Industry Insights.</p>
<h3>1.  Alibaba Launches Qwen-Audio-3.1 With Five Upgraded Models and Major Price Cuts</h3>
<p>Alibaba&#8217;s Qwen team released Qwen-Audio-3.1, upgrading its ASR, TTS, and Realtime audio models and adding two new models, TTS-Next and ASR-Next. The five-model lineup covers audio understanding, generation, interaction, and creation. ASR-Next adds multi-speaker recognition with speaker labels and timestamps plus sound and emotion understanding, while TTS-Next combines a language model with diffusion to generate voice, effects, and background audio in one pass. Alibaba also cut prices across the lineup, with TTS down about 70%, Realtime about 85%, and ASR up to 95% cheaper.<br />
Read more: <a href="https://mp.weixin.qq.com/s/MW6wBdCJEdgGYX1DE9FzXQ">https://mp.weixin.qq.com/s/MW6wBdCJEdgGYX1DE9FzXQ</a></p>
<p><video width="600" height="400" controls poster="https://cdn.ainative.foundation/image/20260929_93f21e82bb74454b851aeabf111b12cc.jpg"><source src="https://cdn.ainative.foundation/video/20260929_gn_video_qwen.mp4" type="video/mp4"></video><br />
Video Credit: The original article</p>
<h3>2.  Kling announces Kling 4.0 video generation model with 4.0 Flash version available for early access</h3>
<p>Kling announced that its upgraded Kling 4.0 video generation model will officially launch in October, with the Kling 4.0 Flash version opening for limited early access on September 28. The 4.0 version delivers major improvements in visual realism, creative control, and narrative capabilities, including enhanced camera movement handling, 4K 10-bit HDR output support, up to 30-second native video generation, multi-keyframe control with up to 10 frames, improved lip-sync accuracy, and support for up to 15 multimodal references including 10 images and 5 videos. The Flash version is designed for high-frequency creation scenarios with faster generation speed and cost efficiency, currently available to premium subscribers.<br />
Read more: <a href="https://mp.weixin.qq.com/s?__biz=MzU5NTkwNDU2OA==&#038;mid=2247496900&#038;idx=1&#038;sn=b0e629ecd3df2f3183f3585eef30a7f1&#038;chksm=ffb9a9cfedc47b5cc650301acfd86535336dbfa10b49e35a848f25113284fbe66c9913db51b9&#038;scene=0&#038;xtrack=1#rd">https://mp.weixin.qq.com/s?__biz=MzU5NTkwNDU2OA==&#038;mid=2247496900&#038;idx=1&#038;sn=b0e629ecd3df2f3183f3585eef30a7f1&#038;chksm=ffb9a9cfedc47b5cc650301acfd86535336dbfa10b49e35a848f25113284fbe66c9913db51b9&#038;scene=0&#038;xtrack=1#rd</a></p>
<p><video width="600" height="400" controls poster="https://cdn.ainative.foundation/image/20260929_c8c1c33f0b4d4470af3eb76e12b49d44"><source src="https://cdn.ainative.foundation/video/20260929_gn_video_kling.mp4" type="video/mp4"></video><br />
Video Credit: The original article</p>
<h3>3.  SkyProduction launches AI MV creation, Mureka music models integration, and intelligent soundtrack features</h3>
<p>On September 27, SkyProduction (Tiangong Workbench) released three new capabilities: AI MV creation tools, integration of Mureka V8, V9, and V9.5 music generation models into its infinite canvas, and intelligent soundtrack generation for videos and images. The AI MV feature analyzes song structure, rhythm, and lyrics to generate director plans and storyboards, then creates videos using the Seedance 2.5 model while maintaining character and scene consistency across shots. The music generation models allow users to create original songs from text prompts, while the intelligent soundtrack feature generates custom background music based on video or image content. All three features operate within a unified canvas workspace, enabling creators to generate music, create music videos, and add soundtracks without switching between different tools.<br />
Read more: <a href="https://mp.weixin.qq.com/s?__biz=MzI1MzE1NDc3Mg==&#038;mid=2247510445&#038;idx=1&#038;sn=986335f6da07fef6af3111ddf76db4bb&#038;chksm=e871f887a0d891d5be2ac5c53b0e391289a46e0990e95944ded8f78c0a4b25588462eae54b8a&#038;scene=0&#038;xtrack=1#rd">https://mp.weixin.qq.com/s?__biz=MzI1MzE1NDc3Mg==&#038;mid=2247510445&#038;idx=1&#038;sn=986335f6da07fef6af3111ddf76db4bb&#038;chksm=e871f887a0d891d5be2ac5c53b0e391289a46e0990e95944ded8f78c0a4b25588462eae54b8a&#038;scene=0&#038;xtrack=1#rd</a></p>
<p><video width="600" height="400" controls poster="https://cdn.ainative.foundation/image/20260929_e8ecf0c6025a4266988af4419e0a744d"><source src="https://cdn.ainative.foundation/video/20260929_gn_video_kunlun.mp4" type="video/mp4"></video><br />
Video Credit: The original article</p>
<div style="width:100%;height:2px;background:#808080;margin:10px 0"></div>
<p>That&#8217;s all for today&#8217;s China AI Native Industry Insights. Join us at <a href="https://member.ainativefoundation.org/">AI Native Foundation Membership Dashboard</a> for the latest insights on AI Native, or follow our linkedin account at <a href="https://www.linkedin.com/company/ainativefoundation/">AI Native Foundation</a> and our twitter account at <a href="https://x.com/AINativeF">AINativeF</a>.</p>
<p>The post <a href="https://ainativefoundation.org/china-ai-native-industry-insights-20260929-alibaba-kling-ai-kunlun-tech-more/">China AI Native Industry Insights &#8211; 20260929 &#8211;  Alibaba | Kling AI | Kunlun Tech | more</a> appeared first on <a href="https://ainativefoundation.org">AI Native Foundation</a>.</p>
]]></content:encoded>
					
		
		<enclosure url="https://cdn.ainative.foundation/video/20260929_gn_video_qwen.mp4" length="63923142" type="video/mp4" />
<enclosure url="https://cdn.ainative.foundation/video/20260929_gn_video_kling.mp4" length="21619604" type="video/mp4" />
<enclosure url="https://cdn.ainative.foundation/video/20260929_gn_video_kunlun.mp4" length="14277965" type="video/mp4" />

			</item>
		<item>
		<title>Global AI Native Industry Insights &#8211; 20260928 &#8211;  Perplexity AI &#124; Nvidia &#124; Google &#124; more</title>
		<link>https://ainativefoundation.org/global-ai-native-industry-insights-20260928-perplexity-ai-nvidia-google-more/</link>
		
		<dc:creator><![CDATA[AINF]]></dc:creator>
		<pubDate>Mon, 28 Sep 2026 08:58:48 +0000</pubDate>
				<category><![CDATA[Global Industry]]></category>
		<guid isPermaLink="false">https://ainativefoundation.org/?p=11020</guid>

					<description><![CDATA[<p>Today's digest features notable advancements from Perplexity and Nvidia in collaboration with Google DeepMind. The broader theme of today's news revolves around enhanced AI capabilities and open scientific collaboration. Perplexity introduces a sophisticated Fast Search feature utilizing the new Rust-based Photon retrieval engine, enhancing search efficiency. Nvidia and Google DeepMind, along with partners, have open-sourced AI-predicted protein complex structures for over 2,800 viruses, advancing the field of virology. Additionally, Google's introduction of Gemma 4 on-device support to the Antigravity SDK highlights an effort to bolster AI's on-device processing power.</p>
<p>The post <a href="https://ainativefoundation.org/global-ai-native-industry-insights-20260928-perplexity-ai-nvidia-google-more/">Global AI Native Industry Insights &#8211; 20260928 &#8211;  Perplexity AI | Nvidia | Google | more</a> appeared first on <a href="https://ainativefoundation.org">AI Native Foundation</a>.</p>
]]></description>
										<content:encoded><![CDATA[<p>Today&#8217;s digest features notable advancements from Perplexity and Nvidia in collaboration with Google DeepMind. The broader theme of today&#8217;s news revolves around enhanced AI capabilities and open scientific collaboration. Perplexity introduces a sophisticated Fast Search feature utilizing the new Rust-based Photon retrieval engine, enhancing search efficiency. Nvidia and Google DeepMind, along with partners, have open-sourced AI-predicted protein complex structures for over 2,800 viruses, advancing the field of virology. Additionally, Google&#8217;s introduction of Gemma 4 on-device support to the Antigravity SDK highlights an effort to bolster AI&#8217;s on-device processing power. Discover more in Today&#8217;s Global AI Native Industry Insights.</p>
<h3>1.  Perplexity launches Fast Search powered by new Rust-based Photon retrieval engine</h3>
<p>Perplexity AI introduced Fast Search, a new feature in its Search API. It runs on Photon, a newly built Rust-based retrieval and ranking service developed by a small team of engineers with the help of hundreds of AI agents. Fast Search returns 95% of search results in 230 milliseconds or less. The release aims to significantly improve search latency for API developers.<br />
Read more: <a href="https://www.perplexity.ai/hub/blog/photon">https://www.perplexity.ai/hub/blog/photon</a></p>
<p><video width="600" height="400" controls poster="https://cdn.ainative.foundation/image/20260928_a462c93569cc4e8aa6d28627c75c0498.png"><source src="https://cdn.ainative.foundation/video/20260928_gj_video_perplexity.mp4" type="video/mp4"></video><br />
Video Credit: NotebookLM</p>
<h3>2.  Nvidia, Google DeepMind and Partners Open Source AI-Predicted Protein Complex Structures for Over 2,800 Viruses</h3>
<p>Nvidia, in collaboration with Google DeepMind, EMBL-EBI, and other research partners, has released AI-predicted protein complex structures for more than 2,800 viruses as an open resource. The dataset aims to give scientists a head start in studying viral biology and preparing for potential disease outbreaks. The structures were generated using AI-based structure prediction methods rather than traditional experimental techniques. Making this data openly available is intended to accelerate research into antiviral treatments and outbreak preparedness across the scientific community.<br />
Read more: <a href="https://blogs.nvidia.com/blog/open-protein-dataset/">https://blogs.nvidia.com/blog/open-protein-dataset/</a></p>
<p><video width="600" height="400" controls poster="https://cdn.ainative.foundation/image/20260928_e244f38446564b3290b8f6247cfcb87f.jpg"><source src="https://cdn.ainative.foundation/video/20260928_6e813fc5f00d49f7974ef9cdb7674ae7.mp4" type="video/mp4"></video><br />
Video Credit: @nvidia on X</p>
<h3>3.  Google Brings Gemma 4 On-Device Support to Antigravity SDK</h3>
<p>Google announced that Gemma 4 can now run locally on-device within the Antigravity SDK. The update lets developers build fully local or hybrid multi-agent workflows that combine cloud models with a local Gemma 4 workforce. Google says the on-device setup, powered by LiteRT, allows code auditing, patching, and testing with full data privacy and no API fees. The feature targets developers seeking cost-efficient and privacy-preserving AI development tools.<br />
Read more: <a href="https://developers.googleblog.com/introducing-support-for-local-ai-models-in-the-antigravity-sdk/">https://developers.googleblog.com/introducing-support-for-local-ai-models-in-the-antigravity-sdk/</a></p>
<p><video width="600" height="400" controls poster="https://cdn.ainative.foundation/image/20260928_439beb41ebef491db166bc50f3c12370.jpg"><source src="https://cdn.ainative.foundation/video/20260928_gj_video_google.mp4" type="video/mp4"></video><br />
Video Credit: @googledevs on X</p>
<div style="width:100%;height:2px;background:#808080;margin:10px 0"></div>
<p>That&#8217;s all for today&#8217;s Global AI Native Industry Insights. Join us at <a href="https://member.ainativefoundation.org/">AI Native Foundation Membership Dashboard</a> for the latest insights on AI Native, or follow our linkedin account at <a href="https://www.linkedin.com/company/ainativefoundation/">AI Native Foundation</a> and our twitter account at <a href="https://x.com/AINativeF">AINativeF</a>.</p>
<p>The post <a href="https://ainativefoundation.org/global-ai-native-industry-insights-20260928-perplexity-ai-nvidia-google-more/">Global AI Native Industry Insights &#8211; 20260928 &#8211;  Perplexity AI | Nvidia | Google | more</a> appeared first on <a href="https://ainativefoundation.org">AI Native Foundation</a>.</p>
]]></content:encoded>
					
		
		<enclosure url="https://cdn.ainative.foundation/video/20260928_gj_video_perplexity.mp4" length="6006100" type="video/mp4" />
<enclosure url="https://cdn.ainative.foundation/video/20260928_6e813fc5f00d49f7974ef9cdb7674ae7.mp4" length="447363" type="video/mp4" />
<enclosure url="https://cdn.ainative.foundation/video/20260928_gj_video_google.mp4" length="102030" type="video/mp4" />

			</item>
		<item>
		<title>China AI Native Industry Insights &#8211; 20260924 &#8211;  PixVerse &#124; Alibaba &#124; Qoder &#124; more</title>
		<link>https://ainativefoundation.org/china-ai-native-industry-insights-20260924-pixverse-alibaba-qoder-more/</link>
		
		<dc:creator><![CDATA[AINF]]></dc:creator>
		<pubDate>Thu, 24 Sep 2026 08:40:26 +0000</pubDate>
				<category><![CDATA[China Industry]]></category>
		<guid isPermaLink="false">https://ainativefoundation.org/?p=11013</guid>

					<description><![CDATA[<p>Today's digest highlights notable advancements from PixVerse and Alibaba's Qwen. The overarching theme is the refinement and diversification of intelligent agent systems. PixVerse has introduced R2, a real-time world model designed to enhance interactive story generation, pushing the limits of dynamic narrative creation. Meanwhile, Qwen Intelligence unveils three state-of-the-art mobile agents, signifying a major leap in versatile AI capabilities. In addition, Qoder Cloud Agents has launched a new Batch feature, optimizing large-scale task execution for agentic operations.</p>
<p>The post <a href="https://ainativefoundation.org/china-ai-native-industry-insights-20260924-pixverse-alibaba-qoder-more/">China AI Native Industry Insights &#8211; 20260924 &#8211;  PixVerse | Alibaba | Qoder | more</a> appeared first on <a href="https://ainativefoundation.org">AI Native Foundation</a>.</p>
]]></description>
										<content:encoded><![CDATA[<p>Today&#8217;s digest highlights notable advancements from PixVerse and Alibaba&#8217;s Qwen. The overarching theme is the refinement and diversification of intelligent agent systems. PixVerse has introduced R2, a real-time world model designed to enhance interactive story generation, pushing the limits of dynamic narrative creation. Meanwhile, Qwen Intelligence unveils three state-of-the-art mobile agents, signifying a major leap in versatile AI capabilities. In addition, Qoder Cloud Agents has launched a new Batch feature, optimizing large-scale task execution for agentic operations. Discover more in Today&#8217;s China AI Native Industry Insights.</p>
<h3>1.  PixVerse launches R2, a real-time world model for interactive story generation</h3>
<p>PixVerse has introduced PixVerse R2, described as a new real-time world model. The system lets users explore generated worlds and control or edit them using text prompts. It supports shaping ongoing storylines and includes characters that can remember context and respond accordingly. The announcement was made via the company&#8217;s official account on September 22, 2026.<br />
Read more: <a href="https://mp.weixin.qq.com/s/vjIPBLJXa_Plcbxidv_Qiw">https://mp.weixin.qq.com/s/vjIPBLJXa_Plcbxidv_Qiw</a></p>
<p><video width="600" height="400" controls poster="https://cdn.ainative.foundation/image/20260924_6a6ae674be4443999862ce7667e4f04c.jpg"><source src="https://cdn.ainative.foundation/video/20260924_aa467fd9bf794d2ba816249fee4b0739.mp4" type="video/mp4"></video><br />
Video Credit: @PixVerse_ on X</p>
<h3>2.  Alibaba&#8217;s Qwen Launches &#8216;Qwen Intelligence&#8217; with Three SOTA Mobile Agents</h3>
<p>Alibaba&#8217;s Qwen team introduced Qwen Intelligence, a personal AI system built around three specialized agents. The Mobile Planner Agent ranks first on MobilePA-Bench and its Business and Memory tracks for task planning and orchestration. The Mobile-Use Agent achieves 82.1 on MobileWorld, 92.2 on MobileWorld-Real, and 97.2 on AndroidDaily, with a 90% end-to-end task success rate, while the Mobile Creative Agent generates images in about 3 seconds, roughly twice as fast as leading competitors. Qwen also open-sourced its benchmark suite, including MobilePA-Bench, MobileWorld, MobileWorld-Real, and MobileWorld-Safety, covering planning, cross-app execution, real-device performance, and safety evaluation.<br />
Read more: <a href="https://www.qwenintelligence.com/">https://www.qwenintelligence.com/</a></p>
<p><video width="600" height="400" controls poster="https://cdn.ainative.foundation/image/20260924_48e7f3df742d4da8bb1dae04f9d03e3d.jpg"><source src="https://cdn.ainative.foundation/video/20260924_gn_video_qwen.mp4" type="video/mp4"></video><br />
Video Credit: The original article</p>
<h3>3.  Qoder Cloud Agents launches Batch feature for large-scale Agent task execution</h3>
<p>Qoder Cloud Agents has released a Batch feature that enables large-scale, asynchronous execution of AI Agent tasks. The feature is designed for scenarios involving thousands of independent tasks that share similar processing logic, such as customer service quality checks, contract reviews, and Agent evaluation testing. Each batch supports up to 10,000 tasks submitted via JSONL files, with built-in state management, failure recovery, and result reconciliation. The platform offers off-peak scheduling between 22:00 and 08:00 Beijing time with promotional discounts on Qwen series models to reduce operational costs.<br />
Read more: <a href="https://mp.weixin.qq.com/s?__biz=MzE5ODI2NzI5Nw==&#038;mid=2247492649&#038;idx=1&#038;sn=535a3a994a1c8422804dfdf83f8eac8f&#038;chksm=976b8e897a195e7ac5c1cf849ae3f549e2caa9144be2f4f023c12680a36c70861325ffa0c9d5&#038;scene=0&#038;xtrack=1#rd">https://mp.weixin.qq.com/s?__biz=MzE5ODI2NzI5Nw==&#038;mid=2247492649&#038;idx=1&#038;sn=535a3a994a1c8422804dfdf83f8eac8f&#038;chksm=976b8e897a195e7ac5c1cf849ae3f549e2caa9144be2f4f023c12680a36c70861325ffa0c9d5&#038;scene=0&#038;xtrack=1#rd</a></p>
<p><video width="600" height="400" controls poster="https://cdn.ainative.foundation/image/20260924_d42a455eeefb40828df4fbef7273cc58"><source src="https://cdn.ainative.foundation/video/20260924_gn_video_qoder.mp4" type="video/mp4"></video><br />
Video Credit: NotebookLM</p>
<div style="width:100%;height:2px;background:#808080;margin:10px 0"></div>
<p>That&#8217;s all for today&#8217;s China AI Native Industry Insights. Join us at <a href="https://member.ainativefoundation.org/">AI Native Foundation Membership Dashboard</a> for the latest insights on AI Native, or follow our linkedin account at <a href="https://www.linkedin.com/company/ainativefoundation/">AI Native Foundation</a> and our twitter account at <a href="https://x.com/AINativeF">AINativeF</a>.</p>
<p>The post <a href="https://ainativefoundation.org/china-ai-native-industry-insights-20260924-pixverse-alibaba-qoder-more/">China AI Native Industry Insights &#8211; 20260924 &#8211;  PixVerse | Alibaba | Qoder | more</a> appeared first on <a href="https://ainativefoundation.org">AI Native Foundation</a>.</p>
]]></content:encoded>
					
		
		<enclosure url="https://cdn.ainative.foundation/video/20260924_aa467fd9bf794d2ba816249fee4b0739.mp4" length="15680417" type="video/mp4" />
<enclosure url="https://cdn.ainative.foundation/video/20260924_gn_video_qwen.mp4" length="2031670" type="video/mp4" />
<enclosure url="https://cdn.ainative.foundation/video/20260924_gn_video_qoder.mp4" length="7109070" type="video/mp4" />

			</item>
		<item>
		<title>Global AI Native Industry Insights &#8211; 20260923 &#8211;  Anthropic &#124; xAI &#124; OpenAI &#124; more</title>
		<link>https://ainativefoundation.org/global-ai-native-industry-insights-20260923-anthropic-xai-openai-more/</link>
		
		<dc:creator><![CDATA[AINF]]></dc:creator>
		<pubDate>Wed, 23 Sep 2026 09:19:29 +0000</pubDate>
				<category><![CDATA[Global Industry]]></category>
		<guid isPermaLink="false">https://ainativefoundation.org/?p=11008</guid>

					<description><![CDATA[<p>Today’s digest spotlights significant advancements by Anthropic, OpenAI, and Google DeepMind. A central theme across these updates is enhanced AI model performance coupled with cost-efficiency. OpenAI's new GPT-6 Sol and GPT-6 Luna models boast a 50% reduction in API pricing, making high-capacity AI more accessible. Meanwhile, Google DeepMind has introduced the Gemini 3.8 Live models, including Extended Thinking Audio capabilities, emphasizing richer and more interactive user experiences. Anthropic's introduction of Claude Opus 5.5 marks the first release in the Claude 5.5 family, underscoring innovation in AI model development.</p>
<p>The post <a href="https://ainativefoundation.org/global-ai-native-industry-insights-20260923-anthropic-xai-openai-more/">Global AI Native Industry Insights &#8211; 20260923 &#8211;  Anthropic | xAI | OpenAI | more</a> appeared first on <a href="https://ainativefoundation.org">AI Native Foundation</a>.</p>
]]></description>
										<content:encoded><![CDATA[<p>Today’s digest spotlights significant advancements by Anthropic, OpenAI, and Google DeepMind. A central theme across these updates is enhanced AI model performance coupled with cost-efficiency. OpenAI&#8217;s new GPT-6 Sol and GPT-6 Luna models boast a 50% reduction in API pricing, making high-capacity AI more accessible. Meanwhile, Google DeepMind has introduced the Gemini 3.8 Live models, including Extended Thinking Audio capabilities, emphasizing richer and more interactive user experiences. Anthropic&#8217;s introduction of Claude Opus 5.5 marks the first release in the Claude 5.5 family, underscoring innovation in AI model development. Discover more in Today&#8217;s Global AI Native Industry Insights.</p>
<h3>1.  Anthropic Launches Claude Opus 5.5, First Model in New Claude 5.5 Family</h3>
<p>Anthropic has introduced Claude Opus 5.5, the first release in its new Claude 5.5 model family. The company says the model matches the performance of Claude Fable 5.1 on most tasks. Opus 5.5 is priced 40% lower to run than its predecessor, Opus 5. The release signals continued cost efficiency gains alongside performance improvements in Anthropic&#8217;s model lineup.<br />
Read more: <a href="https://www.anthropic.com/claude-opus-5-5">https://www.anthropic.com/claude-opus-5-5</a></p>
<p><video width="600" height="400" controls poster="https://cdn.ainative.foundation/image/20260923_70c142ec14b645afaab5430320918b7c.jpg"><source src="https://cdn.ainative.foundation/video/20260923_8dabe84138b24852a413660308f2d24b.mp4" type="video/mp4"></video><br />
Video Credit: @claudeai on X</p>
<h3>2.  xAI releases Grok 4.7 with improved performance at same price and speed</h3>
<p>xAI announced the release of Grok 4.7, a new version of its AI model. The company said Grok 4.7 offers a notable improvement over its predecessor, Grok 4.6. The upgrade is offered at the same price and speed as the previous version. The announcement was made via xAI&#8217;s official social media account.<br />
Read more: <a href="https://x.ai/news/grok-4-7">https://x.ai/news/grok-4-7</a></p>
<p><video width="600" height="400" controls poster="https://cdn.ainative.foundation/image/20260923_d0166bd5ffc1448e85e58471110629c6.png"><source src="https://cdn.ainative.foundation/video/20260923_gj_video_xai.mp4" type="video/mp4"></video><br />
Video Credit: @xai on X</p>
<h3>3.  OpenAI launches GPT-6 Sol and GPT-6 Luna with 50% lower API pricing</h3>
<p>OpenAI introduced two new models, GPT-6 Sol and GPT-6 Luna, expanding the GPT-6 lineup alongside the earlier GPT-6 Astra. The new models carry over much of Astra&#8217;s capabilities while running faster and at lower cost, aimed at supporting large-scale workloads. OpenAI said improvements to caching and inference efficiency allowed it to cut API prices for Sol and Luna by 50% compared with GPT-5.6 promotional pricing.<br />
Read more: <a href="https://openai.com/index/introducing-gpt-6-sol-and-luna/">https://openai.com/index/introducing-gpt-6-sol-and-luna/</a></p>
<p><video width="600" height="400" controls poster="https://cdn.ainative.foundation/image/20260923_107915d939294cd38c89ac070b96cbb3.jpg"><source src="https://cdn.ainative.foundation/video/20260923_d38f1566f5474f95b366b2d344879eee.mp4" type="video/mp4"></video><br />
Video Credit: @OpenAI on X</p>
<h3>4.  Google DeepMind Launches Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking Audio Models</h3>
<p>Google DeepMind introduced two new audio models, Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking. Google&#8217;s team has been building applications with the models spanning creative use cases, real-time customer support, and robotics. The announcement highlights early internal use cases demonstrating the models&#8217; real-time audio capabilities. Further technical details were not specified in the source tweet.<br />
Read more: <a href="https://x.com/i/web/status/2102146240968270047">https://x.com/i/web/status/2102146240968270047</a></p>
<p><video width="600" height="400" controls poster="https://cdn.ainative.foundation/image/20260923_gj_img_google.jpg"><source src="https://cdn.ainative.foundation/video/20260923_gj_video_google.mp4" type="video/mp4"></video><br />
Video Credit: @Google on X</p>
<div style="width:100%;height:2px;background:#808080;margin:10px 0"></div>
<p>That&#8217;s all for today&#8217;s Global AI Native Industry Insights. Join us at <a href="https://member.ainativefoundation.org/">AI Native Foundation Membership Dashboard</a> for the latest insights on AI Native, or follow our linkedin account at <a href="https://www.linkedin.com/company/ainativefoundation/">AI Native Foundation</a> and our twitter account at <a href="https://x.com/AINativeF">AINativeF</a>.</p>
<p>The post <a href="https://ainativefoundation.org/global-ai-native-industry-insights-20260923-anthropic-xai-openai-more/">Global AI Native Industry Insights &#8211; 20260923 &#8211;  Anthropic | xAI | OpenAI | more</a> appeared first on <a href="https://ainativefoundation.org">AI Native Foundation</a>.</p>
]]></content:encoded>
					
		
		<enclosure url="https://cdn.ainative.foundation/video/20260923_8dabe84138b24852a413660308f2d24b.mp4" length="13134325" type="video/mp4" />
<enclosure url="https://cdn.ainative.foundation/video/20260923_gj_video_xai.mp4" length="1510044" type="video/mp4" />
<enclosure url="https://cdn.ainative.foundation/video/20260923_d38f1566f5474f95b366b2d344879eee.mp4" length="1915702" type="video/mp4" />
<enclosure url="https://cdn.ainative.foundation/video/20260923_gj_video_google.mp4" length="26950511" type="video/mp4" />

			</item>
		<item>
		<title>China AI Native Industry Insights &#8211; 20260922 &#8211;  Alibaba &#124; xiaomi &#124; Unitree &#124; more</title>
		<link>https://ainativefoundation.org/china-ai-native-industry-insights-20260922-alibaba-xiaomi-unitree-more/</link>
		
		<dc:creator><![CDATA[AINF]]></dc:creator>
		<pubDate>Tue, 22 Sep 2026 08:07:27 +0000</pubDate>
				<category><![CDATA[China Industry]]></category>
		<guid isPermaLink="false">https://ainativefoundation.org/?p=11003</guid>

					<description><![CDATA[<p>Today's digest spotlights major advancements from Alibaba and Xiaomi, among others. A unifying theme of innovation in image and AI technology runs through these updates, showcasing new benchmarks in AI capabilities. Alibaba's Qwen-Image-2.1 offers open weights for image generation and editing, while Xiaomi's MiMo-V2.6 Pro and Flash Omnimodal models advance multimodal AI applications. Additionally, Tencent's Hy Image 3.5 provides a cost-effective solution for professional creative work, reflecting a trend towards accessible AI-driven tools. Unitree's Dex5-S robotic hand, featuring 22 degrees of freedom, adds to this wave of technological progress.</p>
<p>The post <a href="https://ainativefoundation.org/china-ai-native-industry-insights-20260922-alibaba-xiaomi-unitree-more/">China AI Native Industry Insights &#8211; 20260922 &#8211;  Alibaba | xiaomi | Unitree | more</a> appeared first on <a href="https://ainativefoundation.org">AI Native Foundation</a>.</p>
]]></description>
										<content:encoded><![CDATA[<p>Today&#8217;s digest spotlights major advancements from Alibaba and Xiaomi, among others. A unifying theme of innovation in image and AI technology runs through these updates, showcasing new benchmarks in AI capabilities. Alibaba&#8217;s Qwen-Image-2.1 offers open weights for image generation and editing, while Xiaomi&#8217;s MiMo-V2.6 Pro and Flash Omnimodal models advance multimodal AI applications. Additionally, Tencent&#8217;s Hy Image 3.5 provides a cost-effective solution for professional creative work, reflecting a trend towards accessible AI-driven tools. Unitree&#8217;s Dex5-S robotic hand, featuring 22 degrees of freedom, adds to this wave of technological progress. Discover more in Today&#8217;s China AI Native Industry Insights.</p>
<h3>1.  Alibaba Releases Qwen-Image-2.1 with Open Weights for Unified Image Generation and Editing</h3>
<p>Alibaba&#8217;s Qwen team released Qwen-Image-2.1, a new open-weight model in the Qwen-Image series described as balanced and cost-effective. The model uses a lightweight 7B architecture that unifies image generation and editing, with natively supported RGBA transparency for layered compositing and text editing. It supports up to 10 reference images with high-fidelity local editing for portraits and products, and performs well on panoramas, infographics, and virtual try-ons. Model weights and code are available on Hugging Face, ModelScope, and GitHub.<br />
Read more: <a href="https://www.modelscope.cn/models/Qwen/Qwen-Image-2.1">https://www.modelscope.cn/models/Qwen/Qwen-Image-2.1</a></p>
<p><video width="600" height="400" controls poster="https://cdn.ainative.foundation/image/20260922_ad9d4870be0e478c9ea1ebf89f81ba71.jpg"><source src="https://cdn.ainative.foundation/video/20260922_gn_video_qwen.mp4" type="video/mp4"></video><br />
Video Credit: NotebookLM</p>
<h3>2.  Xiaomi Releases MiMo-V2.6 Pro and Flash Omnimodal AI Models</h3>
<p>Xiaomi introduced MiMo-V2.6, a new family of omnimodal models comprising Pro and Flash versions, developed using scaled reinforcement learning. The company says the Pro model performs on par with Claude Opus 5 and GPT-5.6 Sol across most agent benchmarks and scores 46 on the Artificial Analysis Intelligence Index, the highest among open-source models. Xiaomi highlighted improvements in coding, computer use, 3D reasoning, and creative capabilities. The release includes open model weights, a technical report, RL environments, and training code.<br />
Read more: <a href="https://mimo.xiaomi.com/mimo-v2-6">https://mimo.xiaomi.com/mimo-v2-6</a></p>
<p><video width="600" height="400" controls poster="https://cdn.ainative.foundation/image/20260922_6885c8fe24134311b8728ccec7d12ffa.jpg"><source src="https://cdn.ainative.foundation/video/20260922_gn_video_xaomi.mp4" type="video/mp4"></video><br />
Video Credit: NotebookLM</p>
<h3>3.  Unitree Unveils Dex5-S Dexterous Robotic Hand with 22 Degrees of Freedom</h3>
<p>Unitree Robotics introduced the Dex5-S, a precision biomimetic dexterous hand with 22 degrees of freedom and a size matching a real human hand. All 22 joints support smooth backdrivability and are equipped with limit impact torque protection. The hand is priced starting at $6,500, excluding tax and shipping. The announcement was made via Unitree&#8217;s official social media account.<br />
Read more: <a href="https://mp.weixin.qq.com/s/RTl7NbTGo_IvvrqbaeI7nQ">https://mp.weixin.qq.com/s/RTl7NbTGo_IvvrqbaeI7nQ</a></p>
<p><video width="600" height="400" controls poster="https://cdn.ainative.foundation/image/20260922_feba951606604beca4038c65c207decb.jpg"><source src="https://cdn.ainative.foundation/video/20260922_9925400b024e44a2bcc11a7b890e47ed.mp4" type="video/mp4"></video><br />
Video Credit: @UnitreeRobotics on X</p>
<h3>4.  Tencent releases Hy Image 3.5 preview, a cost-effective image generation model for professional creative work</h3>
<p>Tencent has officially launched Hy Image 3.5 preview, a professional-grade image generation model showing approximately 30% capability improvement over Hy Image 3.0 based on internal blind testing by hundreds of professional designers. The model supports text-to-image and image-to-image generation with up to 5 reference images, multiple aspect ratios, and up to 2K resolution output. It is integrated into multiple Tencent applications including Yuanbao AI assistant, WorkRally, OnSolo, Miora, WorkBuddy, and ima, and is available via Tencent Cloud API at a competitive price of 0.15 yuan per 2K image. The model demonstrates enhanced capabilities in text rendering and layout, realism, style expression, and editing consistency for both everyday creative tasks and professional production scenarios including commercial posters, film production, UI design, and e-commerce materials.<br />
Read more: <a href="https://mp.weixin.qq.com/s?__biz=MzkwODU2OTQyNQ==&#038;mid=2247498632&#038;idx=1&#038;sn=503321c4942df7d8983d81839f1292de&#038;chksm=c1047ee7fcf9efd59761318565e538a470d85cb794b03cef1142ed8332bd951c46349b52c9bf&#038;scene=0&#038;xtrack=1#rd">https://mp.weixin.qq.com/s?__biz=MzkwODU2OTQyNQ==&#038;mid=2247498632&#038;idx=1&#038;sn=503321c4942df7d8983d81839f1292de&#038;chksm=c1047ee7fcf9efd59761318565e538a470d85cb794b03cef1142ed8332bd951c46349b52c9bf&#038;scene=0&#038;xtrack=1#rd</a></p>
<p><video width="600" height="400" controls poster="https://cdn.ainative.foundation/image/20260922_gn_img_hy.png"><source src="https://cdn.ainative.foundation/video/20260922_gn_video_hy.mp4" type="video/mp4"></video><br />
Video Credit: NotebookLM</p>
<div style="width:100%;height:2px;background:#808080;margin:10px 0"></div>
<p>That&#8217;s all for today&#8217;s China AI Native Industry Insights. Join us at <a href="https://member.ainativefoundation.org/">AI Native Foundation Membership Dashboard</a> for the latest insights on AI Native, or follow our linkedin account at <a href="https://www.linkedin.com/company/ainativefoundation/">AI Native Foundation</a> and our twitter account at <a href="https://x.com/AINativeF">AINativeF</a>.</p>
<p>The post <a href="https://ainativefoundation.org/china-ai-native-industry-insights-20260922-alibaba-xiaomi-unitree-more/">China AI Native Industry Insights &#8211; 20260922 &#8211;  Alibaba | xiaomi | Unitree | more</a> appeared first on <a href="https://ainativefoundation.org">AI Native Foundation</a>.</p>
]]></content:encoded>
					
		
		<enclosure url="https://cdn.ainative.foundation/video/20260922_gn_video_qwen.mp4" length="10021173" type="video/mp4" />
<enclosure url="https://cdn.ainative.foundation/video/20260922_gn_video_xaomi.mp4" length="9390503" type="video/mp4" />
<enclosure url="https://cdn.ainative.foundation/video/20260922_9925400b024e44a2bcc11a7b890e47ed.mp4" length="21945516" type="video/mp4" />
<enclosure url="https://cdn.ainative.foundation/video/20260922_gn_video_hy.mp4" length="10246375" type="video/mp4" />

			</item>
	</channel>
</rss>
