{"id":6910,"date":"2026-09-20T17:46:39","date_gmt":"2026-09-20T17:46:39","guid":{"rendered":"https:\/\/areeblog.com\/?p=6910"},"modified":"2026-09-20T17:46:39","modified_gmt":"2026-09-20T17:46:39","slug":"alibabas-new-qwen3-8-omni-flash-pushes-ai-agents-into-long-form-video-and-audio","status":"publish","type":"post","link":"https:\/\/areeblog.com\/alibabas-new-qwen3-8-omni-flash-pushes-ai-agents-into-long-form-video-and-audio\/","title":{"rendered":"Alibaba\u2019s New Qwen3.8-Omni-Flash Pushes AI Agents Into Long-Form Video and Audio"},"content":{"rendered":"<p><img loading=\"lazy\" loading=\"lazy\" decoding=\"async\" data-attachment-id=\"6911\" data-permalink=\"https:\/\/areeblog.com\/alibabas-new-qwen3-8-omni-flash-pushes-ai-agents-into-long-form-video-and-audio\/alibaba-qwen_gxn5nkddxy\/\" data-orig-file=\"https:\/\/areeblog.com\/wp-content\/uploads\/2026\/09\/alibaba-qwen_GXn5NKdDXY.jpg\" data-orig-size=\"840,560\" data-comments-opened=\"1\" data-image-meta=\"{&quot;aperture&quot;:&quot;0&quot;,&quot;credit&quot;:&quot;&quot;,&quot;camera&quot;:&quot;&quot;,&quot;caption&quot;:&quot;&quot;,&quot;created_timestamp&quot;:&quot;0&quot;,&quot;copyright&quot;:&quot;&quot;,&quot;focal_length&quot;:&quot;0&quot;,&quot;iso&quot;:&quot;0&quot;,&quot;shutter_speed&quot;:&quot;0&quot;,&quot;title&quot;:&quot;&quot;,&quot;orientation&quot;:&quot;0&quot;,&quot;alt&quot;:&quot;&quot;}\" data-image-title=\"alibaba-qwen_GXn5NKdDXY\" data-image-description=\"\" data-image-caption=\"\" data-large-file=\"https:\/\/areeblog.com\/wp-content\/uploads\/2026\/09\/alibaba-qwen_GXn5NKdDXY.jpg\" class=\"aligncenter size-full wp-image-6911\" src=\"https:\/\/areeblog.com\/wp-content\/uploads\/2026\/09\/alibaba-qwen_GXn5NKdDXY.jpg\" alt=\"Alibaba\u2019s New Qwen3.8-Omni-Flash Pushes AI Agents Into Long-Form Video and Audio\" width=\"840\" height=\"560\" srcset=\"https:\/\/areeblog.com\/wp-content\/uploads\/2026\/09\/alibaba-qwen_GXn5NKdDXY.jpg 840w, https:\/\/areeblog.com\/wp-content\/uploads\/2026\/09\/alibaba-qwen_GXn5NKdDXY-300x200.jpg 300w, https:\/\/areeblog.com\/wp-content\/uploads\/2026\/09\/alibaba-qwen_GXn5NKdDXY-768x512.jpg 768w, https:\/\/areeblog.com\/wp-content\/uploads\/2026\/09\/alibaba-qwen_GXn5NKdDXY-330x220.jpg 330w, https:\/\/areeblog.com\/wp-content\/uploads\/2026\/09\/alibaba-qwen_GXn5NKdDXY-420x280.jpg 420w, https:\/\/areeblog.com\/wp-content\/uploads\/2026\/09\/alibaba-qwen_GXn5NKdDXY-615x410.jpg 615w\" sizes=\"auto, (max-width: 840px) 100vw, 840px\" \/><\/p>\n<p>Alibaba has introduced Qwen3.8-Omni-Flash, a new native omnimodal <a href=\"https:\/\/areeblog.com\/openai-says-its-ai-models-broke-containment-during-cybersecurity-testing\/\">AI model<\/a> designed to move artificial intelligence agents beyond understanding audio and video and toward planning tasks, using tools and completing work.<\/p>\n<p>The model accepts text, images, audio and video as input and supports a 1-million-token context window. Alibaba says it is designed for real-world productivity workflows including video editing, music video creation, film production and commentary, audiovisual summarization, translation and real-time conversations.<\/p>\n<p>Unlike systems focused mainly on describing what appears in a recording, Qwen3.8-Omni-Flash is designed to determine what a user needs from long-form media and then gather the relevant evidence through multiple stages of analysis.<\/p>\n<p>Alibaba says the model can start from a question, determine what portions of a long video or audio recording need closer examination, retrieve relevant sections and progressively verify information without processing every frame from beginning to end.<\/p>\n<p>On OmniVideoBench, Alibaba reports that its agentic approach increased accuracy from 63.4 to 67.8 while reducing token consumption from 145,736 to 79,117 tokens per query, a reduction of approximately 45.7 percent.<\/p>\n<p>The company says the model improved its average score by more than 25 percent over Qwen3.5-Omni-Plus across 29 evaluations. The evaluation set covers audio reasoning, audiovisual reasoning and audiovisual agent benchmarks.<\/p>\n<p>Alibaba reports a 36.5-point improvement on WildClawBench-MM and a 22.3-point improvement on AgenticVBench, while scoring 69.6 on UniClawBench.<\/p>\n<p>For long-form media, Alibaba reports an 8.3-point gain on LongAudioSpan and a 9.6-point gain on OmniVideoBench. OmniCap-IF CSR improved by 8.5 points and ISR by 14.1 points.<\/p>\n<p>For the company&#8217;s reported AliMeeting evaluation, diarization error rate fell from 88.11 to 3.35, while concatenated minimum-permutation word error rate declined from 89.61 to 17.18.<\/p>\n<p>Alibaba says Qwen3.8-Omni-Flash delivers audiovisual performance close to Gemini 3.8 Flash and overall audio performance that exceeds Gemini 3.8 Flash in its reported evaluations.<\/p>\n<p>On OmniVideoBench, Alibaba reports scores of 63.4 for Qwen3.8-Omni-Flash in static mode and 67.8 when used with Qwen Code. Gemini 3.8 Flash scored 65.2 in static mode and 70.1 with Qwen Code.<\/p>\n<p>On Video-MME-v2, Qwen3.8-Omni-Flash scored 65.0 in static mode and 71.3 with Qwen Code, compared with 71.0 and 72.7 for Gemini 3.8 Flash.<\/p>\n<p>On LVOmniBench, Qwen3.8-Omni-Flash scored 63.3 in static mode and 73.6 with Qwen Code, while Gemini 3.8 Flash scored 70.7 in both settings.<\/p>\n<p>Alibaba says the pricing of the new model has also changed the economics of processing long-form media. The company reports that its API price per hour of audio input is more than 98 percent lower and audio-visual input is more than 93 percent lower than the previous generation.<\/p>\n<p>Alibaba&#8217;s pricing methodology estimates hourly audio or audiovisual input prices as 30 times the input cost of two minutes of source material. Its audiovisual comparison uses 720p video at one frame per second. Alibaba also specifies that Gemini 3.8 Flash was tested with media_resolution set to high, Seed 2.0 Lite with max_frame_tokens set to 384, while other API parameters used their default values.<\/p>\n<p>The model is available through Alibaba Cloud Model Studio in Beijing, Singapore, Hong Kong, Tokyo, Frankfurt and Virginia. Alibaba&#8217;s documentation lists text, image, audio and video inputs with text output, custom function calling, web search and context caching.<\/p>\n<p>The model supports a 1-million-token context window, with a maximum input length of 991,808 tokens in non-thinking mode and 983,616 tokens in thinking mode. The maximum output length is 131,072 tokens.<\/p>\n<p>Alibaba says Qwen3.8-Omni-Flash supports 113 audio languages and dialects and also supports multichannel spatial audio input.<\/p>\n<p>The model supports Chat Completions and Responses APIs. Its reasoning effort can be adjusted, with xhigh as the default setting, while medium and low settings can be used to reduce reasoning depth, latency or cost.<\/p>\n<p>Alibaba is building additional tooling around the model rather than treating the model itself as the complete system.<\/p>\n<p>The company has expanded <a href=\"https:\/\/github.com\/QwenLM\/Qwen-MM-Plugins\">Qwen-MM-Plugins<\/a> with on-demand perception, tool use and workflow execution for long-form audio and video. Alibaba has also open-sourced <a href=\"https:\/\/github.com\/QwenLM\/Qwen-MM-Plugins\">Qwen-Live Harness<\/a> as a runtime for continuous, real-time omnimodal interaction.<\/p>\n<p>For long-form video, Qwen3.8-Omni-Flash can provide different kinds of analysis depending on what the user requests. Alibaba says the same video can be summarized at a high level, searched for specific segments or examined in detail for character actions, camera shots, lighting and sound.<\/p>\n<p>The model can also handle meetings containing multiple participants. Alibaba says it supports up to one hour of audiovisual input for this workflow and can perform speaker segmentation, transcription and identity alignment.<\/p>\n<p>Given a complete meeting video and a request, Alibaba says the system can identify participant relationships, generate meeting minutes, identify action items and analyze project risks. With agents and tool use, it can also send emails, organize tasks and begin coding based on meeting requirements.<\/p>\n<p>Alibaba is also positioning the model for audiovisual research. The company says Qwen3.8-Omni-Flash can combine a user&#8217;s question with video content, identify issues requiring additional investigation, search multimodal sources across the web, including images, videos and documents, and produce a richly illustrated research report.<\/p>\n<p>One example provided by Alibaba involves a Photoshop tutorial showing color fringing around a hair cutout. The model can break down the tutorial, examine the principles behind Multiply and Screen blend modes, compare alternative edge-repair methods and explain which approach best fits the situation.<\/p>\n<p>For music video production, Alibaba says Qwen3.8-Omni-Flash can analyze a song&#8217;s structure, rhythm, mood, vocals and instrumental changes. It can use that analysis to support decisions about characters, scenes and shots and can produce line-level lyrics with timestamps for synchronizing vocals, subtitles and visuals.<\/p>\n<p>Through Qwen-MM-Plugins, Alibaba says the music video workflow can extend from music analysis and creative planning to production and final quality review.<\/p>\n<p>The company has also demonstrated a short-drama translation workflow designed to combine speaker-aware dialogue recognition, conversational translation, character voice cloning, dubbing, audio remixing and final quality review.<\/p>\n<p>Alibaba says an agent built around Qwen3.8-Omni-Flash can handle those stages from a single natural-language instruction rather than requiring users to coordinate separate transcription, translation, dubbing and editing systems manually.<\/p>\n<p>For long-form film commentary, Alibaba says the system can process two- or three-hour films, extract key plot points, plan commentary, produce voiceover and music, edit and render the result, and perform a final quality review.<\/p>\n<p>The system can also interleave original dialogue with commentary and adjust speech rate and volume so narration, original audio, background music and visuals work together.<\/p>\n<p>Alibaba is extending the model&#8217;s use beyond content production into model development itself.<\/p>\n<p>In one experiment, Qwen3.8-Omni-Flash was tasked with improving Sichuan dialect speech recognition for Qwen2.5-Omni-3B within 12 hours. Alibaba says the agent selected the WenetSpeech-Chuan evaluation set, established evaluation criteria and a baseline, listened to audio samples, diagnosed recognition problems and generated targeted training data.<\/p>\n<p>Across four rounds of experiments, the agent created 3,413 training examples, changed its approach based on evaluation results, kept effective improvements and rolled back unsuccessful attempts.<\/p>\n<p>Alibaba says the character error rate of Qwen2.5-Omni-3B on the same evaluation set fell from 25.79 percent to 15.30 percent, a relative reduction of approximately 40.7 percent.<\/p>\n<p>The company is also using Qwen3.8-Omni-Flash to turn long-form audiovisual material into more compact information resources.<\/p>\n<p>Alibaba has open-sourced <a href=\"https:\/\/github.com\/QwenLM\/Qwen-MM-Plugins\">Video2Note<\/a>, which uses the model&#8217;s understanding of speech, visuals and procedures to organize knowledge, break down key steps, select representative video frames and generate PDF notes containing text and images.<\/p>\n<p>Alibaba says automated review and iterative correction can turn hours of instructional video into structured documents that can be revisited more easily.<\/p>\n<p>Another open-source capability, <a href=\"https:\/\/github.com\/QwenLM\/Qwen-MM-Plugins\">Omni Skill Creator<\/a>, is designed to extract standard operating procedures from demonstrations and capture tool usage, decision criteria and practical knowledge from expert instruction.<\/p>\n<p>Alibaba says a single demonstration can be transformed into a reusable agent skill that is verified and evaluated and can then be shared for automated work.<\/p>\n<p>Qwen3.8-Omni-Flash also has a separate real-time counterpart called Qwen3.8-Omni-Flash-Realtime. Alibaba says the model is designed to receive live audiovisual streams, respond with low latency, maintain real-time context, call tools and execute tasks.<\/p>\n<p>Alibaba describes real-time speaking practice as one application. The system jointly models pronunciation and semantics, handles expressions affected by accents and can provide standard-pronunciation demonstrations during practice.<\/p>\n<p>The company also says its realtime system can combine spatial sound with visual information to determine the direction and distance of sound sources while accounting for obstacles, navigable areas and changes in the environment.<\/p>\n<p>Alibaba says this allows instructions such as locating a sound to be connected to tool-based localization, search, path planning and navigation.<\/p>\n<p>The realtime model can also load identity settings, expression styles, business knowledge and interaction rules through Skills. Alibaba says this can allow the same system to adapt to customer-service requirements while processing speech, visuals and contextual information.<\/p>\n<p>Alibaba&#8217;s documentation for the realtime system lists 74 speech-recognition languages and 39 Chinese dialects, while speech generation supports 29 languages and seven Chinese dialects.<\/p>\n<p>Qwen3.8-Omni-Flash therefore represents a broader shift in Alibaba&#8217;s Qwen strategy. The company is treating audio and video not simply as additional inputs for multimodal models, but as environments in which agents can locate evidence, reason across time, call external tools and execute multi-step workflows.<\/p>\n<p>The model itself produces text output in its standard configuration. Its audiovisual production capabilities depend on the surrounding agent tools and workflows rather than making Qwen3.8-Omni-Flash a standalone video-generation model.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Alibaba has introduced Qwen3.8-Omni-Flash, a new native omnimodal AI model designed to move artificial intelligence agents beyond understanding audio and video and toward planning tasks, using tools and completing work. The model accepts text, images, audio and video as input and supports a 1-million-token context window. Alibaba says it is designed for real-world productivity workflows [&hellip;]<\/p>\n","protected":false},"author":2,"featured_media":6911,"comment_status":"open","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"content-type":"","_monsterinsights_skip_tracking":false,"_jetpack_newsletter_access":"","_jetpack_dont_email_post_to_subs":false,"_jetpack_newsletter_tier_id":0,"_jetpack_memberships_contains_paywalled_content":false,"_jetpack_feature_clip_id":0,"_jetpack_memberships_contains_paid_content":false,"footnotes":"","jetpack_post_was_ever_published":false},"categories":[164],"tags":[166],"class_list":["post-6910","post","type-post","status-publish","format-standard","has-post-thumbnail","category-tech-updates","tag-ai"],"share_on_mastodon":{"url":"https:\/\/mastodon.social\/@Areeblog\/117304638948211255","error":""},"yoast_head":"<!-- This site is optimized with the Yoast SEO Premium plugin v28.4 (Yoast SEO v28.5) - https:\/\/yoast.com\/product\/yoast-seo-premium-wordpress\/ -->\n<title>Alibaba\u2019s New Qwen3.8-Omni-Flash Pushes AI Agents Into Long-Form Video and Audio - Aree Blog<\/title>\n<meta name=\"description\" content=\"Alibaba\u2019s Qwen3.8-Omni-Flash brings AI agents to long-form video and audio with 1M-token context and tool use.\" \/>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/areeblog.com\/alibabas-new-qwen3-8-omni-flash-pushes-ai-agents-into-long-form-video-and-audio\/\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"Alibaba\u2019s New Qwen3.8-Omni-Flash Pushes AI Agents Into Long-Form Video and Audio\" \/>\n<meta property=\"og:description\" content=\"Alibaba\u2019s Qwen3.8-Omni-Flash brings AI agents to long-form video and audio with 1M-token context and tool use.\" \/>\n<meta property=\"og:url\" content=\"https:\/\/areeblog.com\/alibabas-new-qwen3-8-omni-flash-pushes-ai-agents-into-long-form-video-and-audio\/\" \/>\n<meta property=\"og:site_name\" content=\"Aree Blog\" \/>\n<meta property=\"article:published_time\" content=\"2026-09-20T17:46:39+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/areeblog.com\/wp-content\/uploads\/2026\/09\/alibaba-qwen_GXn5NKdDXY.jpg\" \/>\n\t<meta property=\"og:image:width\" content=\"840\" \/>\n\t<meta property=\"og:image:height\" content=\"560\" \/>\n\t<meta property=\"og:image:type\" content=\"image\/jpeg\" \/>\n<meta name=\"author\" content=\"Daniel Chinonso John\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"Daniel Chinonso John\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"8 minutes\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\\\/\\\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\\\/\\\/areeblog.com\\\/alibabas-new-qwen3-8-omni-flash-pushes-ai-agents-into-long-form-video-and-audio\\\/#article\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/areeblog.com\\\/alibabas-new-qwen3-8-omni-flash-pushes-ai-agents-into-long-form-video-and-audio\\\/\"},\"author\":{\"name\":\"Daniel Chinonso John\",\"@id\":\"https:\\\/\\\/areeblog.com\\\/#\\\/schema\\\/person\\\/d972222c55618fb0f4b4c0c11ff52f63\"},\"headline\":\"Alibaba\u2019s New Qwen3.8-Omni-Flash Pushes AI Agents Into Long-Form Video and Audio\",\"datePublished\":\"2026-09-20T17:46:39+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\\\/\\\/areeblog.com\\\/alibabas-new-qwen3-8-omni-flash-pushes-ai-agents-into-long-form-video-and-audio\\\/\"},\"wordCount\":1530,\"commentCount\":0,\"image\":{\"@id\":\"https:\\\/\\\/areeblog.com\\\/alibabas-new-qwen3-8-omni-flash-pushes-ai-agents-into-long-form-video-and-audio\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/areeblog.com\\\/wp-content\\\/uploads\\\/2026\\\/09\\\/alibaba-qwen_GXn5NKdDXY.jpg\",\"keywords\":[\"AI\"],\"articleSection\":[\"Tech Updates\"],\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"CommentAction\",\"name\":\"Comment\",\"target\":[\"https:\\\/\\\/areeblog.com\\\/alibabas-new-qwen3-8-omni-flash-pushes-ai-agents-into-long-form-video-and-audio\\\/#respond\"]}]},{\"@type\":\"WebPage\",\"@id\":\"https:\\\/\\\/areeblog.com\\\/alibabas-new-qwen3-8-omni-flash-pushes-ai-agents-into-long-form-video-and-audio\\\/\",\"url\":\"https:\\\/\\\/areeblog.com\\\/alibabas-new-qwen3-8-omni-flash-pushes-ai-agents-into-long-form-video-and-audio\\\/\",\"name\":\"Alibaba\u2019s New Qwen3.8-Omni-Flash Pushes AI Agents Into Long-Form Video and Audio - Aree Blog\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/areeblog.com\\\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\\\/\\\/areeblog.com\\\/alibabas-new-qwen3-8-omni-flash-pushes-ai-agents-into-long-form-video-and-audio\\\/#primaryimage\"},\"image\":{\"@id\":\"https:\\\/\\\/areeblog.com\\\/alibabas-new-qwen3-8-omni-flash-pushes-ai-agents-into-long-form-video-and-audio\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/areeblog.com\\\/wp-content\\\/uploads\\\/2026\\\/09\\\/alibaba-qwen_GXn5NKdDXY.jpg\",\"datePublished\":\"2026-09-20T17:46:39+00:00\",\"author\":{\"@id\":\"https:\\\/\\\/areeblog.com\\\/#\\\/schema\\\/person\\\/d972222c55618fb0f4b4c0c11ff52f63\"},\"description\":\"Alibaba\u2019s Qwen3.8-Omni-Flash brings AI agents to long-form video and audio with 1M-token context and tool use.\",\"breadcrumb\":{\"@id\":\"https:\\\/\\\/areeblog.com\\\/alibabas-new-qwen3-8-omni-flash-pushes-ai-agents-into-long-form-video-and-audio\\\/#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\\\/\\\/areeblog.com\\\/alibabas-new-qwen3-8-omni-flash-pushes-ai-agents-into-long-form-video-and-audio\\\/\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/areeblog.com\\\/alibabas-new-qwen3-8-omni-flash-pushes-ai-agents-into-long-form-video-and-audio\\\/#primaryimage\",\"url\":\"https:\\\/\\\/areeblog.com\\\/wp-content\\\/uploads\\\/2026\\\/09\\\/alibaba-qwen_GXn5NKdDXY.jpg\",\"contentUrl\":\"https:\\\/\\\/areeblog.com\\\/wp-content\\\/uploads\\\/2026\\\/09\\\/alibaba-qwen_GXn5NKdDXY.jpg\",\"width\":840,\"height\":560,\"caption\":\"Alibaba\u2019s New Qwen3.8-Omni-Flash Pushes AI Agents Into Long-Form Video and Audio\"},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\\\/\\\/areeblog.com\\\/alibabas-new-qwen3-8-omni-flash-pushes-ai-agents-into-long-form-video-and-audio\\\/#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\\\/\\\/areeblog.com\\\/\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"Alibaba\u2019s New Qwen3.8-Omni-Flash Pushes AI Agents Into Long-Form Video and Audio\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\\\/\\\/areeblog.com\\\/#website\",\"url\":\"https:\\\/\\\/areeblog.com\\\/\",\"name\":\"Aree Blog\",\"description\":\"Unfiltered Perspectives, Unstoppable Insights\",\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\\\/\\\/areeblog.com\\\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"},{\"@type\":\"Person\",\"@id\":\"https:\\\/\\\/areeblog.com\\\/#\\\/schema\\\/person\\\/d972222c55618fb0f4b4c0c11ff52f63\",\"name\":\"Daniel Chinonso John\",\"description\":\"Daniel Chinonso John is a web designer, penetration tester, and founder of Aree Tech. He writes clear, actionable posts at the intersection of productivity, AI, cybersecurity, and blogging to help readers get things done.\",\"sameAs\":[\"https:\\\/\\\/www.linkedin.com\\\/in\\\/daniel-john-45183a169\\\/\"],\"url\":\"https:\\\/\\\/areeblog.com\\\/author\\\/danojohn55gmail-com\\\/\"}]}<\/script>\n<!-- \/ Yoast SEO Premium plugin. -->","yoast_head_json":{"title":"Alibaba\u2019s New Qwen3.8-Omni-Flash Pushes AI Agents Into Long-Form Video and Audio - Aree Blog","description":"Alibaba\u2019s Qwen3.8-Omni-Flash brings AI agents to long-form video and audio with 1M-token context and tool use.","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/areeblog.com\/alibabas-new-qwen3-8-omni-flash-pushes-ai-agents-into-long-form-video-and-audio\/","og_locale":"en_US","og_type":"article","og_title":"Alibaba\u2019s New Qwen3.8-Omni-Flash Pushes AI Agents Into Long-Form Video and Audio","og_description":"Alibaba\u2019s Qwen3.8-Omni-Flash brings AI agents to long-form video and audio with 1M-token context and tool use.","og_url":"https:\/\/areeblog.com\/alibabas-new-qwen3-8-omni-flash-pushes-ai-agents-into-long-form-video-and-audio\/","og_site_name":"Aree Blog","article_published_time":"2026-09-20T17:46:39+00:00","og_image":[{"width":840,"height":560,"url":"https:\/\/areeblog.com\/wp-content\/uploads\/2026\/09\/alibaba-qwen_GXn5NKdDXY.jpg","type":"image\/jpeg"}],"author":"Daniel Chinonso John","twitter_card":"summary_large_image","twitter_misc":{"Written by":"Daniel Chinonso John","Est. reading time":"8 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/areeblog.com\/alibabas-new-qwen3-8-omni-flash-pushes-ai-agents-into-long-form-video-and-audio\/#article","isPartOf":{"@id":"https:\/\/areeblog.com\/alibabas-new-qwen3-8-omni-flash-pushes-ai-agents-into-long-form-video-and-audio\/"},"author":{"name":"Daniel Chinonso John","@id":"https:\/\/areeblog.com\/#\/schema\/person\/d972222c55618fb0f4b4c0c11ff52f63"},"headline":"Alibaba\u2019s New Qwen3.8-Omni-Flash Pushes AI Agents Into Long-Form Video and Audio","datePublished":"2026-09-20T17:46:39+00:00","mainEntityOfPage":{"@id":"https:\/\/areeblog.com\/alibabas-new-qwen3-8-omni-flash-pushes-ai-agents-into-long-form-video-and-audio\/"},"wordCount":1530,"commentCount":0,"image":{"@id":"https:\/\/areeblog.com\/alibabas-new-qwen3-8-omni-flash-pushes-ai-agents-into-long-form-video-and-audio\/#primaryimage"},"thumbnailUrl":"https:\/\/areeblog.com\/wp-content\/uploads\/2026\/09\/alibaba-qwen_GXn5NKdDXY.jpg","keywords":["AI"],"articleSection":["Tech Updates"],"inLanguage":"en-US","potentialAction":[{"@type":"CommentAction","name":"Comment","target":["https:\/\/areeblog.com\/alibabas-new-qwen3-8-omni-flash-pushes-ai-agents-into-long-form-video-and-audio\/#respond"]}]},{"@type":"WebPage","@id":"https:\/\/areeblog.com\/alibabas-new-qwen3-8-omni-flash-pushes-ai-agents-into-long-form-video-and-audio\/","url":"https:\/\/areeblog.com\/alibabas-new-qwen3-8-omni-flash-pushes-ai-agents-into-long-form-video-and-audio\/","name":"Alibaba\u2019s New Qwen3.8-Omni-Flash Pushes AI Agents Into Long-Form Video and Audio - Aree Blog","isPartOf":{"@id":"https:\/\/areeblog.com\/#website"},"primaryImageOfPage":{"@id":"https:\/\/areeblog.com\/alibabas-new-qwen3-8-omni-flash-pushes-ai-agents-into-long-form-video-and-audio\/#primaryimage"},"image":{"@id":"https:\/\/areeblog.com\/alibabas-new-qwen3-8-omni-flash-pushes-ai-agents-into-long-form-video-and-audio\/#primaryimage"},"thumbnailUrl":"https:\/\/areeblog.com\/wp-content\/uploads\/2026\/09\/alibaba-qwen_GXn5NKdDXY.jpg","datePublished":"2026-09-20T17:46:39+00:00","author":{"@id":"https:\/\/areeblog.com\/#\/schema\/person\/d972222c55618fb0f4b4c0c11ff52f63"},"description":"Alibaba\u2019s Qwen3.8-Omni-Flash brings AI agents to long-form video and audio with 1M-token context and tool use.","breadcrumb":{"@id":"https:\/\/areeblog.com\/alibabas-new-qwen3-8-omni-flash-pushes-ai-agents-into-long-form-video-and-audio\/#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/areeblog.com\/alibabas-new-qwen3-8-omni-flash-pushes-ai-agents-into-long-form-video-and-audio\/"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/areeblog.com\/alibabas-new-qwen3-8-omni-flash-pushes-ai-agents-into-long-form-video-and-audio\/#primaryimage","url":"https:\/\/areeblog.com\/wp-content\/uploads\/2026\/09\/alibaba-qwen_GXn5NKdDXY.jpg","contentUrl":"https:\/\/areeblog.com\/wp-content\/uploads\/2026\/09\/alibaba-qwen_GXn5NKdDXY.jpg","width":840,"height":560,"caption":"Alibaba\u2019s New Qwen3.8-Omni-Flash Pushes AI Agents Into Long-Form Video and Audio"},{"@type":"BreadcrumbList","@id":"https:\/\/areeblog.com\/alibabas-new-qwen3-8-omni-flash-pushes-ai-agents-into-long-form-video-and-audio\/#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/areeblog.com\/"},{"@type":"ListItem","position":2,"name":"Alibaba\u2019s New Qwen3.8-Omni-Flash Pushes AI Agents Into Long-Form Video and Audio"}]},{"@type":"WebSite","@id":"https:\/\/areeblog.com\/#website","url":"https:\/\/areeblog.com\/","name":"Aree Blog","description":"Unfiltered Perspectives, Unstoppable Insights","potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/areeblog.com\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":"Person","@id":"https:\/\/areeblog.com\/#\/schema\/person\/d972222c55618fb0f4b4c0c11ff52f63","name":"Daniel Chinonso John","description":"Daniel Chinonso John is a web designer, penetration tester, and founder of Aree Tech. He writes clear, actionable posts at the intersection of productivity, AI, cybersecurity, and blogging to help readers get things done.","sameAs":["https:\/\/www.linkedin.com\/in\/daniel-john-45183a169\/"],"url":"https:\/\/areeblog.com\/author\/danojohn55gmail-com\/"}]}},"jetpack_sharing_enabled":true,"jetpack-related-posts":[{"id":5265,"url":"https:\/\/areeblog.com\/agents-payments-protocol\/","url_meta":{"origin":6910,"position":0},"title":"Google and Industry Partners Launch Agents Payments Protocol to Power AI Commerce","author":"Samuel Ogori","date":"September 16, 2025","format":false,"excerpt":"Google, working alongside more than 60 payments and technology companies, has rolled out the Agents Payments Protocol (AP2), \u00a0an open standard that allows AI agents to securely complete purchases on behalf of users. This launch is one of the first big steps to bring agent-driven commerce into mainstream payments. It\u2026","rel":"","context":"In &quot;Tech Updates&quot;","block_context":{"text":"Tech Updates","link":"https:\/\/areeblog.com\/category\/tech-updates\/"},"img":{"alt_text":"Google and Industry Partners Launch Agents Payments Protocol to Power AI Commerce","src":"https:\/\/i0.wp.com\/areeblog.com\/wp-content\/uploads\/2025\/09\/Agents-Payments-Protocol.jpg?resize=350%2C200&ssl=1","width":350,"height":200,"srcset":"https:\/\/i0.wp.com\/areeblog.com\/wp-content\/uploads\/2025\/09\/Agents-Payments-Protocol.jpg?resize=350%2C200&ssl=1 1x, https:\/\/i0.wp.com\/areeblog.com\/wp-content\/uploads\/2025\/09\/Agents-Payments-Protocol.jpg?resize=525%2C300&ssl=1 1.5x, https:\/\/i0.wp.com\/areeblog.com\/wp-content\/uploads\/2025\/09\/Agents-Payments-Protocol.jpg?resize=700%2C400&ssl=1 2x, https:\/\/i0.wp.com\/areeblog.com\/wp-content\/uploads\/2025\/09\/Agents-Payments-Protocol.jpg?resize=1050%2C600&ssl=1 3x"},"classes":[]},{"id":6839,"url":"https:\/\/areeblog.com\/google-says-hackers-are-using-ai-agents-to-run-multi-stage-attacks-with-little-human-input\/","url_meta":{"origin":6910,"position":1},"title":"Google Says Hackers Are Using AI Agents to Run Multi-Stage Attacks With Little Human Input","author":"Daniel Chinonso John","date":"September 9, 2026","format":false,"excerpt":"Hackers are increasingly using artificial intelligence to automate multiple stages of cyberattacks, with Google Threat Intelligence Group reporting that some attackers have moved beyond simple prompting to AI-driven workflows capable of scanning targets, troubleshooting failures and harvesting credentials with limited human involvement. In a report published September 8, 2026, Google\u2026","rel":"","context":"In &quot;Tech Updates&quot;","block_context":{"text":"Tech Updates","link":"https:\/\/areeblog.com\/category\/tech-updates\/"},"img":{"alt_text":"Google Says Hackers Are Using AI Agents to Run Multi-Stage Attacks With Little Human Input","src":"https:\/\/i0.wp.com\/areeblog.com\/wp-content\/uploads\/2026\/09\/images-55.jpeg?resize=350%2C200&ssl=1","width":350,"height":200,"srcset":"https:\/\/i0.wp.com\/areeblog.com\/wp-content\/uploads\/2026\/09\/images-55.jpeg?resize=350%2C200&ssl=1 1x, https:\/\/i0.wp.com\/areeblog.com\/wp-content\/uploads\/2026\/09\/images-55.jpeg?resize=525%2C300&ssl=1 1.5x"},"classes":[]},{"id":6835,"url":"https:\/\/areeblog.com\/nsa-fbi-and-cisa-accuse-china-ai-firms-of-stealing-u-s-model-capabilities\/","url_meta":{"origin":6910,"position":2},"title":"NSA, FBI and CISA Accuse China AI Firms of Stealing U.S. Model Capabilities","author":"Daniel Chinonso John","date":"September 8, 2026","format":false,"excerpt":"The U.S. National Security Agency, Federal Bureau of Investigation and Cybersecurity and Infrastructure Security Agency have accused six China-based artificial intelligence companies of conducting industrial-scale campaigns to extract capabilities from U.S. frontier AI models. In a joint cybersecurity advisory released on September 8, 2026, the agencies said the campaigns involve\u2026","rel":"","context":"In &quot;Tech Updates&quot;","block_context":{"text":"Tech Updates","link":"https:\/\/areeblog.com\/category\/tech-updates\/"},"img":{"alt_text":"NSA, FBI and CISA Accuse China AI Firms of Stealing U.S. Model Capabilities","src":"https:\/\/i0.wp.com\/areeblog.com\/wp-content\/uploads\/2026\/09\/970a355d-2615-4e71-9d32-c5dad14b8324_860e4f66.jpg?resize=350%2C200&ssl=1","width":350,"height":200,"srcset":"https:\/\/i0.wp.com\/areeblog.com\/wp-content\/uploads\/2026\/09\/970a355d-2615-4e71-9d32-c5dad14b8324_860e4f66.jpg?resize=350%2C200&ssl=1 1x, https:\/\/i0.wp.com\/areeblog.com\/wp-content\/uploads\/2026\/09\/970a355d-2615-4e71-9d32-c5dad14b8324_860e4f66.jpg?resize=525%2C300&ssl=1 1.5x, https:\/\/i0.wp.com\/areeblog.com\/wp-content\/uploads\/2026\/09\/970a355d-2615-4e71-9d32-c5dad14b8324_860e4f66.jpg?resize=700%2C400&ssl=1 2x, https:\/\/i0.wp.com\/areeblog.com\/wp-content\/uploads\/2026\/09\/970a355d-2615-4e71-9d32-c5dad14b8324_860e4f66.jpg?resize=1050%2C600&ssl=1 3x, https:\/\/i0.wp.com\/areeblog.com\/wp-content\/uploads\/2026\/09\/970a355d-2615-4e71-9d32-c5dad14b8324_860e4f66.jpg?resize=1400%2C800&ssl=1 4x"},"classes":[]},{"id":5205,"url":"https:\/\/areeblog.com\/generative-and-agentic-ai-for-content-personalization-and-automation\/","url_meta":{"origin":6910,"position":3},"title":"Generative and Agentic AI for Content, Personalization and Automation","author":"Samuel Ogori","date":"September 12, 2025","format":false,"excerpt":"Generative AI and agentic AI are moving from experiment to everyday tools for teams that build content, run campaigns, and automate customer journeys. One side creates new assets (text, images, audio, video) at speed. The other side plans, decides, and acts across systems without waiting for a human to push\u2026","rel":"","context":"In &quot;Artificial Intelligence&quot;","block_context":{"text":"Artificial Intelligence","link":"https:\/\/areeblog.com\/category\/artificial-intelligence\/"},"img":{"alt_text":"Generative & Agentic AI for Content, Personalization and Automation","src":"https:\/\/i0.wp.com\/areeblog.com\/wp-content\/uploads\/2025\/09\/Generative-Agentic-AI-Aree-Blog.jpg?resize=350%2C200&ssl=1","width":350,"height":200,"srcset":"https:\/\/i0.wp.com\/areeblog.com\/wp-content\/uploads\/2025\/09\/Generative-Agentic-AI-Aree-Blog.jpg?resize=350%2C200&ssl=1 1x, https:\/\/i0.wp.com\/areeblog.com\/wp-content\/uploads\/2025\/09\/Generative-Agentic-AI-Aree-Blog.jpg?resize=525%2C300&ssl=1 1.5x, https:\/\/i0.wp.com\/areeblog.com\/wp-content\/uploads\/2025\/09\/Generative-Agentic-AI-Aree-Blog.jpg?resize=700%2C400&ssl=1 2x, https:\/\/i0.wp.com\/areeblog.com\/wp-content\/uploads\/2025\/09\/Generative-Agentic-AI-Aree-Blog.jpg?resize=1050%2C600&ssl=1 3x"},"classes":[]},{"id":4905,"url":"https:\/\/areeblog.com\/z-ai-releases-glm-4-5-series-high-performance-open-source-models-focused-on-agent-capabilities\/","url_meta":{"origin":6910,"position":4},"title":"Z.AI Releases GLM 4.5 Series: High-Performance, Open-Source Models Focused on Agent Capabilities","author":"Samuel Ogori","date":"August 2, 2025","format":false,"excerpt":"Z.AI (formerly Zepoo AI) has launched the GLM 4.5 series, comprising the flagship GLM 4.5 model and the lighter GLM 4.5 Air. Positioned as a significant open-source release in 2025, these models emphasize a balance of performance, efficiency, agent capabilities, and cost. Model Architecture and Efficiency GLM 4.5: A 355\u2026","rel":"","context":"In &quot;Tech Updates&quot;","block_context":{"text":"Tech Updates","link":"https:\/\/areeblog.com\/category\/tech-updates\/"},"img":{"alt_text":"Z.AI Releases GLM 4.5 Series: High-Performance, Open-Source Models Focused on Agent Capabilities","src":"https:\/\/i0.wp.com\/areeblog.com\/wp-content\/uploads\/2025\/08\/file_00000000f38c6246b5d683f4a05174d4.png?resize=350%2C200&ssl=1","width":350,"height":200,"srcset":"https:\/\/i0.wp.com\/areeblog.com\/wp-content\/uploads\/2025\/08\/file_00000000f38c6246b5d683f4a05174d4.png?resize=350%2C200&ssl=1 1x, https:\/\/i0.wp.com\/areeblog.com\/wp-content\/uploads\/2025\/08\/file_00000000f38c6246b5d683f4a05174d4.png?resize=525%2C300&ssl=1 1.5x, https:\/\/i0.wp.com\/areeblog.com\/wp-content\/uploads\/2025\/08\/file_00000000f38c6246b5d683f4a05174d4.png?resize=700%2C400&ssl=1 2x"},"classes":[]},{"id":4968,"url":"https:\/\/areeblog.com\/google-notebook-lm-video-overviews\/","url_meta":{"origin":6910,"position":5},"title":"Google Notebook LM Adds Video Overviews","author":"Daniel Chinonso John","date":"August 7, 2025","format":false,"excerpt":"What if you could turn texts into a clear, engaging video you can watch in minutes? That\u2019s exactly what Google\u2019s new video overviews in Notebook LM deliver. In this post, we\u2019ll walk through what Google Notebook LM is, how it works, and why it can reshape the way you learn,\u2026","rel":"","context":"In &quot;Tech Updates&quot;","block_context":{"text":"Tech Updates","link":"https:\/\/areeblog.com\/category\/tech-updates\/"},"img":{"alt_text":"Google Notebook LM Video Adds Overviews","src":"https:\/\/i0.wp.com\/areeblog.com\/wp-content\/uploads\/2025\/08\/Google-Notebook-LM-Video-Adds-Overviews.jpg?resize=350%2C200&ssl=1","width":350,"height":200,"srcset":"https:\/\/i0.wp.com\/areeblog.com\/wp-content\/uploads\/2025\/08\/Google-Notebook-LM-Video-Adds-Overviews.jpg?resize=350%2C200&ssl=1 1x, https:\/\/i0.wp.com\/areeblog.com\/wp-content\/uploads\/2025\/08\/Google-Notebook-LM-Video-Adds-Overviews.jpg?resize=525%2C300&ssl=1 1.5x, https:\/\/i0.wp.com\/areeblog.com\/wp-content\/uploads\/2025\/08\/Google-Notebook-LM-Video-Adds-Overviews.jpg?resize=700%2C400&ssl=1 2x, https:\/\/i0.wp.com\/areeblog.com\/wp-content\/uploads\/2025\/08\/Google-Notebook-LM-Video-Adds-Overviews.jpg?resize=1050%2C600&ssl=1 3x"},"classes":[]}],"jetpack_featured_media_url":"https:\/\/areeblog.com\/wp-content\/uploads\/2026\/09\/alibaba-qwen_GXn5NKdDXY.jpg","_links":{"self":[{"href":"https:\/\/areeblog.com\/wp-json\/wp\/v2\/posts\/6910","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/areeblog.com\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/areeblog.com\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/areeblog.com\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/areeblog.com\/wp-json\/wp\/v2\/comments?post=6910"}],"version-history":[{"count":1,"href":"https:\/\/areeblog.com\/wp-json\/wp\/v2\/posts\/6910\/revisions"}],"predecessor-version":[{"id":6912,"href":"https:\/\/areeblog.com\/wp-json\/wp\/v2\/posts\/6910\/revisions\/6912"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/areeblog.com\/wp-json\/wp\/v2\/media\/6911"}],"wp:attachment":[{"href":"https:\/\/areeblog.com\/wp-json\/wp\/v2\/media?parent=6910"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/areeblog.com\/wp-json\/wp\/v2\/categories?post=6910"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/areeblog.com\/wp-json\/wp\/v2\/tags?post=6910"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}