{"id":6938,"date":"2026-09-23T18:54:18","date_gmt":"2026-09-23T18:54:18","guid":{"rendered":"https:\/\/areeblog.com\/?p=6938"},"modified":"2026-09-23T18:54:18","modified_gmt":"2026-09-23T18:54:18","slug":"nvidia-releases-a-100m-parameter-open-model-that-can-track-eight-speakers-in-real-time","status":"publish","type":"post","link":"https:\/\/areeblog.com\/nvidia-releases-a-100m-parameter-open-model-that-can-track-eight-speakers-in-real-time\/","title":{"rendered":"NVIDIA Releases a 100M-Parameter Open Model That Can Track Eight Speakers in Real Time"},"content":{"rendered":"<p><img loading=\"lazy\" loading=\"lazy\" decoding=\"async\" data-attachment-id=\"6939\" data-permalink=\"https:\/\/areeblog.com\/nvidia-releases-a-100m-parameter-open-model-that-can-track-eight-speakers-in-real-time\/2026-07-16t115806z-1705182164-rc2yemamcyvv-rtrmadp-3-nvidia-japan-2026-07-b91c7076ec607a9fe7bf6bce896062e2\/\" data-orig-file=\"https:\/\/areeblog.com\/wp-content\/uploads\/2026\/09\/2026-07-16t115806z-1705182164-rc2yemamcyvv-rtrmadp-3-nvidia-japan-2026-07-b91c7076ec607a9fe7bf6bce896062e2.jpg\" data-orig-size=\"400,225\" data-comments-opened=\"1\" data-image-meta=\"{&quot;aperture&quot;:&quot;0&quot;,&quot;credit&quot;:&quot;&quot;,&quot;camera&quot;:&quot;&quot;,&quot;caption&quot;:&quot;&quot;,&quot;created_timestamp&quot;:&quot;0&quot;,&quot;copyright&quot;:&quot;&quot;,&quot;focal_length&quot;:&quot;0&quot;,&quot;iso&quot;:&quot;0&quot;,&quot;shutter_speed&quot;:&quot;0&quot;,&quot;title&quot;:&quot;&quot;,&quot;orientation&quot;:&quot;0&quot;,&quot;alt&quot;:&quot;&quot;}\" data-image-title=\"2026-07-16t115806z-1705182164-rc2yemamcyvv-rtrmadp-3-nvidia-japan-2026-07-b91c7076ec607a9fe7bf6bce896062e2\" data-image-description=\"\" data-image-caption=\"\" data-large-file=\"https:\/\/areeblog.com\/wp-content\/uploads\/2026\/09\/2026-07-16t115806z-1705182164-rc2yemamcyvv-rtrmadp-3-nvidia-japan-2026-07-b91c7076ec607a9fe7bf6bce896062e2.jpg\" class=\"aligncenter size-full wp-image-6939\" src=\"https:\/\/areeblog.com\/wp-content\/uploads\/2026\/09\/2026-07-16t115806z-1705182164-rc2yemamcyvv-rtrmadp-3-nvidia-japan-2026-07-b91c7076ec607a9fe7bf6bce896062e2.jpg\" alt=\"NVIDIA Releases a 100M-Parameter Open Model That Can Track Eight Speakers in Real Time\" width=\"400\" height=\"225\" srcset=\"https:\/\/areeblog.com\/wp-content\/uploads\/2026\/09\/2026-07-16t115806z-1705182164-rc2yemamcyvv-rtrmadp-3-nvidia-japan-2026-07-b91c7076ec607a9fe7bf6bce896062e2.jpg 400w, https:\/\/areeblog.com\/wp-content\/uploads\/2026\/09\/2026-07-16t115806z-1705182164-rc2yemamcyvv-rtrmadp-3-nvidia-japan-2026-07-b91c7076ec607a9fe7bf6bce896062e2-300x169.jpg 300w\" sizes=\"auto, (max-width: 400px) 100vw, 400px\" \/><\/p>\n<p><a href=\"https:\/\/areeblog.com\/doj-investigates-nvidias-17-billion-groq-deal-over-potential-antitrust-concerns\/\">NVIDIA<\/a> has released Nemotron 3 Diarization, a 99.2-million-parameter open-weight speech model designed to track up to eight speakers in real time across live and recorded conversations.<\/p>\n<p>The model was released on September 23, 2026, through NVIDIA&#8217;s <a href=\"https:\/\/huggingface.co\/nvidia\/Nemotron-3-Diarization\">Nemotron 3 Diarization model repository on Hugging Face<\/a>. NVIDIA describes it as a roughly 100-million-parameter model, while the published model contains 99.2 million parameters.<\/p>\n<p>Nemotron 3 Diarization is designed for speaker diarization, the process of determining who spoke and when during an audio recording. It does not identify people by name.<\/p>\n<p>Instead, the model produces anonymous speaker channels. An application can assign the first detected speaker to one channel, the next distinct speaker to another, and continue the process for up to eight channels.<\/p>\n<p>The model supports both streaming and offline inference. NVIDIA designed the streaming system to maintain speaker assignments as additional audio arrives rather than treating every incoming segment as an independent recording.<\/p>\n<p>Nemotron 3 accepts 16 kHz mono audio in WAV, FLAC, OPUS and MP3 formats. Its output is a speaker-activity probability tensor with a shape of <code>[T, 8]<\/code>, where the eight output channels correspond to the model&#8217;s maximum supported speaker streams.<\/p>\n<p>The default temporal resolution is 10 milliseconds. NVIDIA says the model technically supports an 80-millisecond minimum buffer, but recommends a 0.32-second buffer as the minimum practical streaming configuration.<\/p>\n<p>The published operating points include 0.32 seconds, 0.64 seconds and 1.04 seconds for streaming, as well as a 30.4-second offline-style configuration.<\/p>\n<p>Those figures describe input-buffer latency rather than complete end-to-end response time. NVIDIA says the measurements exclude model computation, networking, automatic speech recognition and other application processing.<\/p>\n<p>The model&#8217;s eight speaker channels represent an expansion over NVIDIA&#8217;s previous Streaming Sortformer model, which was designed for four speakers. The new system is intended to handle larger group conversations without increasing the number of separate diarization systems required to cover the conversation.<\/p>\n<p>NVIDIA&#8217;s architecture is based on a 31-layer Transformer encoder with Rotary Positional Embeddings. Audio is converted into 10-millisecond Mel-spectrogram frames, stacked by a factor of eight, and processed as 80-millisecond encoder frames before the system produces 10-millisecond speaker activity outputs.<\/p>\n<p>The streaming design includes an Arrival-Order Speaker Cache, or AOSC, together with a first-in-first-out context queue. The cache retains information about speakers encountered earlier in a conversation, while the context queue provides recent frame history to the current inference step.<\/p>\n<p>NVIDIA&#8217;s approach assigns speaker channels according to arrival order. The first distinct voice encountered receives the first channel, followed by subsequent voices. This avoids having to repeatedly solve the speaker-label permutation problem when processing separate streaming chunks.<\/p>\n<p>The model also uses right context, allowing the system to look slightly ahead in incoming audio before making an output decision.<\/p>\n<p>Nemotron 3 is a diarization model, not an automatic speech recognition model. To produce a transcript showing which person said each sentence, developers must combine the diarization output with an ASR system.<\/p>\n<p>NVIDIA&#8217;s documentation gives an example of combining diarization with offline transcription by using word timestamps and a midpoint-based method to assign words to speakers. NVIDIA notes that this is an illustrative alignment approach and does not completely solve cases where multiple people speak at the same time.<\/p>\n<p>The release is aimed at applications including meetings, telephone and customer-service calls, podcasts and speech-recognition systems that require speaker-attributed transcripts.<\/p>\n<p>The speaker-activity signal can also be used for voice activity detection, speaker-specific endpointing, interruption detection and turn-taking in voice-agent systems.<\/p>\n<p>NVIDIA says Nemotron 3 ranked first in the initial VoiceArena Diarization-Bench evaluation, recording a 14.72 percent Diarization Error Rate. The next-ranked system recorded 19.3 percent.<\/p>\n<p>NVIDIA calculates that result as approximately a 24 percent relative reduction in error compared with the next-ranked system.<\/p>\n<p>The initial VoiceArena evaluation covered 139 English-language conversations representing about 22 hours of audio. The comparison included 12 systems and 17 system configurations, with overlapping speech included in the evaluation.<\/p>\n<p>The benchmark used system-generated speech activity detection and a zero-second boundary collar. NVIDIA says the VoiceArena leaderboard may change as the Version 1 evaluation and paired statistical analysis are completed.<\/p>\n<p>Diarization Error Rate combines missed speech, false-alarm time and speaker-confusion errors. The measure does not mean that a 14.72 percent DER is equivalent to a model correctly identifying 85.28 percent of people.<\/p>\n<p>NVIDIA also reports results from a broader internal evaluation covering 901 condition-specific recordings. The evaluation included telephone conversations, meetings, near-field and far-field microphone conditions and multi-microphone environments, with overlapping speech included.<\/p>\n<p>At the 1.04-second latency configuration, Nemotron 3 recorded a 13.18 percent DER on DIHARD III, compared with 19.60 percent for the previous Streaming Sortformer.<\/p>\n<p>On CALLHOME-Part2, Nemotron 3 recorded 10.29 percent compared with 11.31 percent for the earlier model.<\/p>\n<p>On AliMeeting Near, the new model recorded 6.59 percent compared with 12.47 percent. On AliMeeting Far, it recorded 10.80 percent compared with 15.58 percent.<\/p>\n<p>On AMI MHM, Nemotron 3 recorded 9.48 percent compared with 16.36 percent, while AMI SDM recorded 12.80 percent compared with 21.73 percent.<\/p>\n<p>On NOTSOFAR1 MHM, the new model recorded 7.70 percent compared with 22.12 percent. On NOTSOFAR1 SC, it recorded 12.77 percent compared with 31.81 percent.<\/p>\n<p>NVIDIA calculates an unweighted average relative DER reduction of 41.0 percent across those eight conditions at the 1.04-second configuration. That figure is an average of relative reductions rather than a pooled DER across all recordings.<\/p>\n<p>The results are not uniformly better in every condition. NVIDIA gives the two-speaker CALLHOME subset as an example. In that subset at the 30.4-second setting, the earlier model recorded 5.68 percent DER while Nemotron 3 recorded 5.98 percent.<\/p>\n<p>The largest reported gains appear in higher-speaker-count conditions. On NOTSOFAR1 MHM at the 30.4-second setting, the previous model recorded 11.14 percent DER for recordings with three to four speakers and 29.38 percent for recordings with five to seven speakers.<\/p>\n<p>Nemotron 3 recorded 5.25 percent for the three-to-four-speaker group and 7.86 percent for the five-to-seven-speaker group.<\/p>\n<p>The model&#8217;s maximum of eight speakers is a hard limit. NVIDIA warns that recordings with more than eight speakers can result in missed or incorrectly assigned speech.<\/p>\n<p>That limitation is relevant to the DIHARD III evaluation because that benchmark includes recordings with five through nine speakers. A model with eight output channels cannot represent more than eight simultaneous speaker identities within its channel structure.<\/p>\n<p>The model was trained with about 10,000 hours of real conversational audio and 82,611 hours of simulated multi-talker mixtures created from about 28,000 hours of single-speaker recordings.<\/p>\n<p>NVIDIA lists training sources including Fisher, AMI, ICSI, VoxConverse, AISHELL-4, DIHARD III, CALLHOME, AliMeeting, DiPCo, NOTSOFAR1, DISPLACE 2024, DISPLACE-M 2026, licensed David AI multispeaker recordings and a pseudo-labelled YODAS-v2 subset.<\/p>\n<p>The simulated mixtures were created from material including LibriSpeech, AMI, AliMeeting, Fisher and licensed David AI datasets.<\/p>\n<p>NVIDIA also reports that licensed David AI conversational data contributed measurable improvements. Adding the David AI data reduced compound DER from 11.19 percent to 10.42 percent, an absolute reduction of 0.77 percentage points at both the offline-style and ultra-low-latency operating points.<\/p>\n<p>The model card identifies a David AI D12 Human Transcripts dataset covering 21 languages. That does not mean the model&#8217;s published benchmark establishes equal performance across all 21 languages. The detailed evaluation results reported by NVIDIA are centred on the datasets and language conditions described in its testing.<\/p>\n<p>Training began from a Transformer-based NEST self-supervised checkpoint. NVIDIA trained the model using eight nodes containing eight NVIDIA A100-SXM4-80GB GPUs each, for a total of 64 A100 GPUs.<\/p>\n<p>The training process included an offline stage based on simulated data followed by streaming fine-tuning using both real conversations and simulated multi-speaker mixtures.<\/p>\n<p>For inference, NVIDIA reports measurements on an RTX PRO 5000 Blackwell GPU using BF16 precision and the PyTorch backend. The tests compared standard eager execution with <code>torch.compile()<\/code>.<\/p>\n<p>At the 1.04-second latency configuration and batch size one, NVIDIA reports 38 times real-time factor for eager execution and 164 times for the compiled configuration.<\/p>\n<p>At batch size 32, the reported figures increase to 581 times and 865 times real-time factor respectively.<\/p>\n<p>At the 0.32-second configuration, batch size one reached 12.5 times real-time factor with eager execution and 54 times with compilation. At batch size 32, the figures were 199 times and 292 times.<\/p>\n<p>These are hardware- and configuration-specific measurements and should not be treated as universal performance figures for every supported NVIDIA GPU.<\/p>\n<p>The model card lists compatibility across NVIDIA&#8217;s Ampere, Ada Lovelace, Hopper and Blackwell architectures. The documented hardware range includes GeForce GPUs, RTX workstation products and data-centre GPUs such as the A100, H100 and H200.<\/p>\n<p>NVIDIA recommends Linux for deployment.<\/p>\n<p>The current model is released under the <a href=\"https:\/\/openmdw.ai\/license\/1-1\/\">OpenMDW License Agreement 1.1<\/a>. The model repository describes Nemotron 3 Diarization as an open-weight release.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>NVIDIA has released Nemotron 3 Diarization, a 99.2-million-parameter open-weight speech model designed to track up to eight speakers in real time across live and recorded conversations. The model was released on September 23, 2026, through NVIDIA&#8217;s Nemotron 3 Diarization model repository on Hugging Face. NVIDIA describes it as a roughly 100-million-parameter model, while the published [&hellip;]<\/p>\n","protected":false},"author":2,"featured_media":6939,"comment_status":"open","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"content-type":"","_monsterinsights_skip_tracking":false,"_jetpack_newsletter_access":"","_jetpack_dont_email_post_to_subs":false,"_jetpack_newsletter_tier_id":0,"_jetpack_memberships_contains_paywalled_content":false,"_jetpack_feature_clip_id":0,"_jetpack_memberships_contains_paid_content":false,"footnotes":"","jetpack_post_was_ever_published":false},"categories":[164],"tags":[166],"class_list":["post-6938","post","type-post","status-publish","format-standard","has-post-thumbnail","category-tech-updates","tag-ai"],"share_on_mastodon":{"url":"https:\/\/mastodon.social\/@Areeblog\/117321889637329529","error":""},"yoast_head":"<!-- This site is optimized with the Yoast SEO Premium plugin v28.4 (Yoast SEO v28.5) - https:\/\/yoast.com\/product\/yoast-seo-premium-wordpress\/ -->\n<title>NVIDIA Releases a 100M-Parameter Open Model That Can Track Eight Speakers in Real Time - Aree Blog<\/title>\n<meta name=\"description\" content=\"NVIDIA releases a 99.2M-parameter open model that track up to eight speakers in real time for live conversations.\" \/>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/areeblog.com\/nvidia-releases-a-100m-parameter-open-model-that-can-track-eight-speakers-in-real-time\/\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"NVIDIA Releases a 100M-Parameter Open Model That Can Track Eight Speakers in Real Time\" \/>\n<meta property=\"og:description\" content=\"NVIDIA releases a 99.2M-parameter open model that track up to eight speakers in real time for live conversations.\" \/>\n<meta property=\"og:url\" content=\"https:\/\/areeblog.com\/nvidia-releases-a-100m-parameter-open-model-that-can-track-eight-speakers-in-real-time\/\" \/>\n<meta property=\"og:site_name\" content=\"Aree Blog\" \/>\n<meta property=\"article:published_time\" content=\"2026-09-23T18:54:18+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/areeblog.com\/wp-content\/uploads\/2026\/09\/2026-07-16t115806z-1705182164-rc2yemamcyvv-rtrmadp-3-nvidia-japan-2026-07-b91c7076ec607a9fe7bf6bce896062e2.jpg\" \/>\n\t<meta property=\"og:image:width\" content=\"400\" \/>\n\t<meta property=\"og:image:height\" content=\"225\" \/>\n\t<meta property=\"og:image:type\" content=\"image\/jpeg\" \/>\n<meta name=\"author\" content=\"Daniel Chinonso John\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"Daniel Chinonso John\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"7 minutes\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\\\/\\\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\\\/\\\/areeblog.com\\\/nvidia-releases-a-100m-parameter-open-model-that-can-track-eight-speakers-in-real-time\\\/#article\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/areeblog.com\\\/nvidia-releases-a-100m-parameter-open-model-that-can-track-eight-speakers-in-real-time\\\/\"},\"author\":{\"name\":\"Daniel Chinonso John\",\"@id\":\"https:\\\/\\\/areeblog.com\\\/#\\\/schema\\\/person\\\/d972222c55618fb0f4b4c0c11ff52f63\"},\"headline\":\"NVIDIA Releases a 100M-Parameter Open Model That Can Track Eight Speakers in Real Time\",\"datePublished\":\"2026-09-23T18:54:18+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\\\/\\\/areeblog.com\\\/nvidia-releases-a-100m-parameter-open-model-that-can-track-eight-speakers-in-real-time\\\/\"},\"wordCount\":1376,\"commentCount\":0,\"image\":{\"@id\":\"https:\\\/\\\/areeblog.com\\\/nvidia-releases-a-100m-parameter-open-model-that-can-track-eight-speakers-in-real-time\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/areeblog.com\\\/wp-content\\\/uploads\\\/2026\\\/09\\\/2026-07-16t115806z-1705182164-rc2yemamcyvv-rtrmadp-3-nvidia-japan-2026-07-b91c7076ec607a9fe7bf6bce896062e2.jpg\",\"keywords\":[\"AI\"],\"articleSection\":[\"Tech Updates\"],\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"CommentAction\",\"name\":\"Comment\",\"target\":[\"https:\\\/\\\/areeblog.com\\\/nvidia-releases-a-100m-parameter-open-model-that-can-track-eight-speakers-in-real-time\\\/#respond\"]}]},{\"@type\":\"WebPage\",\"@id\":\"https:\\\/\\\/areeblog.com\\\/nvidia-releases-a-100m-parameter-open-model-that-can-track-eight-speakers-in-real-time\\\/\",\"url\":\"https:\\\/\\\/areeblog.com\\\/nvidia-releases-a-100m-parameter-open-model-that-can-track-eight-speakers-in-real-time\\\/\",\"name\":\"NVIDIA Releases a 100M-Parameter Open Model That Can Track Eight Speakers in Real Time - Aree Blog\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/areeblog.com\\\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\\\/\\\/areeblog.com\\\/nvidia-releases-a-100m-parameter-open-model-that-can-track-eight-speakers-in-real-time\\\/#primaryimage\"},\"image\":{\"@id\":\"https:\\\/\\\/areeblog.com\\\/nvidia-releases-a-100m-parameter-open-model-that-can-track-eight-speakers-in-real-time\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/areeblog.com\\\/wp-content\\\/uploads\\\/2026\\\/09\\\/2026-07-16t115806z-1705182164-rc2yemamcyvv-rtrmadp-3-nvidia-japan-2026-07-b91c7076ec607a9fe7bf6bce896062e2.jpg\",\"datePublished\":\"2026-09-23T18:54:18+00:00\",\"author\":{\"@id\":\"https:\\\/\\\/areeblog.com\\\/#\\\/schema\\\/person\\\/d972222c55618fb0f4b4c0c11ff52f63\"},\"description\":\"NVIDIA releases a 99.2M-parameter open model that track up to eight speakers in real time for live conversations.\",\"breadcrumb\":{\"@id\":\"https:\\\/\\\/areeblog.com\\\/nvidia-releases-a-100m-parameter-open-model-that-can-track-eight-speakers-in-real-time\\\/#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\\\/\\\/areeblog.com\\\/nvidia-releases-a-100m-parameter-open-model-that-can-track-eight-speakers-in-real-time\\\/\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/areeblog.com\\\/nvidia-releases-a-100m-parameter-open-model-that-can-track-eight-speakers-in-real-time\\\/#primaryimage\",\"url\":\"https:\\\/\\\/areeblog.com\\\/wp-content\\\/uploads\\\/2026\\\/09\\\/2026-07-16t115806z-1705182164-rc2yemamcyvv-rtrmadp-3-nvidia-japan-2026-07-b91c7076ec607a9fe7bf6bce896062e2.jpg\",\"contentUrl\":\"https:\\\/\\\/areeblog.com\\\/wp-content\\\/uploads\\\/2026\\\/09\\\/2026-07-16t115806z-1705182164-rc2yemamcyvv-rtrmadp-3-nvidia-japan-2026-07-b91c7076ec607a9fe7bf6bce896062e2.jpg\",\"width\":400,\"height\":225,\"caption\":\"NVIDIA Releases a 100M-Parameter Open Model That Can Track Eight Speakers in Real Time\"},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\\\/\\\/areeblog.com\\\/nvidia-releases-a-100m-parameter-open-model-that-can-track-eight-speakers-in-real-time\\\/#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\\\/\\\/areeblog.com\\\/\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"NVIDIA Releases a 100M-Parameter Open Model That Can Track Eight Speakers in Real Time\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\\\/\\\/areeblog.com\\\/#website\",\"url\":\"https:\\\/\\\/areeblog.com\\\/\",\"name\":\"Aree Blog\",\"description\":\"Unfiltered Perspectives, Unstoppable Insights\",\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\\\/\\\/areeblog.com\\\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"},{\"@type\":\"Person\",\"@id\":\"https:\\\/\\\/areeblog.com\\\/#\\\/schema\\\/person\\\/d972222c55618fb0f4b4c0c11ff52f63\",\"name\":\"Daniel Chinonso John\",\"description\":\"Daniel Chinonso John is a web designer, penetration tester, and founder of Aree Tech. He writes clear, actionable posts at the intersection of productivity, AI, cybersecurity, and blogging to help readers get things done.\",\"sameAs\":[\"https:\\\/\\\/www.linkedin.com\\\/in\\\/daniel-john-45183a169\\\/\"],\"url\":\"https:\\\/\\\/areeblog.com\\\/author\\\/danojohn55gmail-com\\\/\"}]}<\/script>\n<!-- \/ Yoast SEO Premium plugin. -->","yoast_head_json":{"title":"NVIDIA Releases a 100M-Parameter Open Model That Can Track Eight Speakers in Real Time - Aree Blog","description":"NVIDIA releases a 99.2M-parameter open model that track up to eight speakers in real time for live conversations.","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/areeblog.com\/nvidia-releases-a-100m-parameter-open-model-that-can-track-eight-speakers-in-real-time\/","og_locale":"en_US","og_type":"article","og_title":"NVIDIA Releases a 100M-Parameter Open Model That Can Track Eight Speakers in Real Time","og_description":"NVIDIA releases a 99.2M-parameter open model that track up to eight speakers in real time for live conversations.","og_url":"https:\/\/areeblog.com\/nvidia-releases-a-100m-parameter-open-model-that-can-track-eight-speakers-in-real-time\/","og_site_name":"Aree Blog","article_published_time":"2026-09-23T18:54:18+00:00","og_image":[{"width":400,"height":225,"url":"https:\/\/areeblog.com\/wp-content\/uploads\/2026\/09\/2026-07-16t115806z-1705182164-rc2yemamcyvv-rtrmadp-3-nvidia-japan-2026-07-b91c7076ec607a9fe7bf6bce896062e2.jpg","type":"image\/jpeg"}],"author":"Daniel Chinonso John","twitter_card":"summary_large_image","twitter_misc":{"Written by":"Daniel Chinonso John","Est. reading time":"7 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/areeblog.com\/nvidia-releases-a-100m-parameter-open-model-that-can-track-eight-speakers-in-real-time\/#article","isPartOf":{"@id":"https:\/\/areeblog.com\/nvidia-releases-a-100m-parameter-open-model-that-can-track-eight-speakers-in-real-time\/"},"author":{"name":"Daniel Chinonso John","@id":"https:\/\/areeblog.com\/#\/schema\/person\/d972222c55618fb0f4b4c0c11ff52f63"},"headline":"NVIDIA Releases a 100M-Parameter Open Model That Can Track Eight Speakers in Real Time","datePublished":"2026-09-23T18:54:18+00:00","mainEntityOfPage":{"@id":"https:\/\/areeblog.com\/nvidia-releases-a-100m-parameter-open-model-that-can-track-eight-speakers-in-real-time\/"},"wordCount":1376,"commentCount":0,"image":{"@id":"https:\/\/areeblog.com\/nvidia-releases-a-100m-parameter-open-model-that-can-track-eight-speakers-in-real-time\/#primaryimage"},"thumbnailUrl":"https:\/\/areeblog.com\/wp-content\/uploads\/2026\/09\/2026-07-16t115806z-1705182164-rc2yemamcyvv-rtrmadp-3-nvidia-japan-2026-07-b91c7076ec607a9fe7bf6bce896062e2.jpg","keywords":["AI"],"articleSection":["Tech Updates"],"inLanguage":"en-US","potentialAction":[{"@type":"CommentAction","name":"Comment","target":["https:\/\/areeblog.com\/nvidia-releases-a-100m-parameter-open-model-that-can-track-eight-speakers-in-real-time\/#respond"]}]},{"@type":"WebPage","@id":"https:\/\/areeblog.com\/nvidia-releases-a-100m-parameter-open-model-that-can-track-eight-speakers-in-real-time\/","url":"https:\/\/areeblog.com\/nvidia-releases-a-100m-parameter-open-model-that-can-track-eight-speakers-in-real-time\/","name":"NVIDIA Releases a 100M-Parameter Open Model That Can Track Eight Speakers in Real Time - Aree Blog","isPartOf":{"@id":"https:\/\/areeblog.com\/#website"},"primaryImageOfPage":{"@id":"https:\/\/areeblog.com\/nvidia-releases-a-100m-parameter-open-model-that-can-track-eight-speakers-in-real-time\/#primaryimage"},"image":{"@id":"https:\/\/areeblog.com\/nvidia-releases-a-100m-parameter-open-model-that-can-track-eight-speakers-in-real-time\/#primaryimage"},"thumbnailUrl":"https:\/\/areeblog.com\/wp-content\/uploads\/2026\/09\/2026-07-16t115806z-1705182164-rc2yemamcyvv-rtrmadp-3-nvidia-japan-2026-07-b91c7076ec607a9fe7bf6bce896062e2.jpg","datePublished":"2026-09-23T18:54:18+00:00","author":{"@id":"https:\/\/areeblog.com\/#\/schema\/person\/d972222c55618fb0f4b4c0c11ff52f63"},"description":"NVIDIA releases a 99.2M-parameter open model that track up to eight speakers in real time for live conversations.","breadcrumb":{"@id":"https:\/\/areeblog.com\/nvidia-releases-a-100m-parameter-open-model-that-can-track-eight-speakers-in-real-time\/#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/areeblog.com\/nvidia-releases-a-100m-parameter-open-model-that-can-track-eight-speakers-in-real-time\/"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/areeblog.com\/nvidia-releases-a-100m-parameter-open-model-that-can-track-eight-speakers-in-real-time\/#primaryimage","url":"https:\/\/areeblog.com\/wp-content\/uploads\/2026\/09\/2026-07-16t115806z-1705182164-rc2yemamcyvv-rtrmadp-3-nvidia-japan-2026-07-b91c7076ec607a9fe7bf6bce896062e2.jpg","contentUrl":"https:\/\/areeblog.com\/wp-content\/uploads\/2026\/09\/2026-07-16t115806z-1705182164-rc2yemamcyvv-rtrmadp-3-nvidia-japan-2026-07-b91c7076ec607a9fe7bf6bce896062e2.jpg","width":400,"height":225,"caption":"NVIDIA Releases a 100M-Parameter Open Model That Can Track Eight Speakers in Real Time"},{"@type":"BreadcrumbList","@id":"https:\/\/areeblog.com\/nvidia-releases-a-100m-parameter-open-model-that-can-track-eight-speakers-in-real-time\/#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/areeblog.com\/"},{"@type":"ListItem","position":2,"name":"NVIDIA Releases a 100M-Parameter Open Model That Can Track Eight Speakers in Real Time"}]},{"@type":"WebSite","@id":"https:\/\/areeblog.com\/#website","url":"https:\/\/areeblog.com\/","name":"Aree Blog","description":"Unfiltered Perspectives, Unstoppable Insights","potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/areeblog.com\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":"Person","@id":"https:\/\/areeblog.com\/#\/schema\/person\/d972222c55618fb0f4b4c0c11ff52f63","name":"Daniel Chinonso John","description":"Daniel Chinonso John is a web designer, penetration tester, and founder of Aree Tech. He writes clear, actionable posts at the intersection of productivity, AI, cybersecurity, and blogging to help readers get things done.","sameAs":["https:\/\/www.linkedin.com\/in\/daniel-john-45183a169\/"],"url":"https:\/\/areeblog.com\/author\/danojohn55gmail-com\/"}]}},"jetpack_sharing_enabled":true,"jetpack-related-posts":[{"id":6349,"url":"https:\/\/areeblog.com\/japan-plans-a-major-rubin-chip-buildout-for-domestic-physical-ai\/","url_meta":{"origin":6938,"position":0},"title":"Japan Plans a Major Rubin Chip Buildout for Domestic Physical AI","author":"Daniel Chinonso John","date":"July 19, 2026","format":false,"excerpt":"Japan is moving ahead with a large AI infrastructure project built around Nvidia\u2019s next-generation Rubin chips. Nvidia said it is working with Noetra Corp. on an AI factory that will use 27,500 Rubin GPUs and 13,750 Vera CPUs, with 140 megawatts of data-center capacity, to support Japan\u2019s FRONTia project. The\u2026","rel":"","context":"In &quot;Tech Updates&quot;","block_context":{"text":"Tech Updates","link":"https:\/\/areeblog.com\/category\/tech-updates\/"},"img":{"alt_text":"Japan Plans a Major Rubin Chip Buildout for Domestic Physical AI","src":"https:\/\/i0.wp.com\/areeblog.com\/wp-content\/uploads\/2026\/07\/images-29.jpeg?resize=350%2C200&ssl=1","width":350,"height":200,"srcset":"https:\/\/i0.wp.com\/areeblog.com\/wp-content\/uploads\/2026\/07\/images-29.jpeg?resize=350%2C200&ssl=1 1x, https:\/\/i0.wp.com\/areeblog.com\/wp-content\/uploads\/2026\/07\/images-29.jpeg?resize=525%2C300&ssl=1 1.5x, https:\/\/i0.wp.com\/areeblog.com\/wp-content\/uploads\/2026\/07\/images-29.jpeg?resize=700%2C400&ssl=1 2x"},"classes":[]},{"id":6456,"url":"https:\/\/areeblog.com\/amds-helios-ai-rack-marks-a-new-phase-in-the-ai-infrastructure-race\/","url_meta":{"origin":6938,"position":1},"title":"AMD\u2019s Helios AI Rack Marks a New Phase in the AI Infrastructure Race","author":"Daniel Chinonso John","date":"August 6, 2026","format":false,"excerpt":"Advanced Micro Devices (AMD) has formally introduced Helios, a rack-scale AI infrastructure platform that reflects how competition in artificial intelligence is expanding beyond standalone graphics processors. Rather than selling GPUs as individual components, the company is packaging compute, networking and software into a single integrated system designed for large-scale AI\u2026","rel":"","context":"In &quot;Tech Updates&quot;","block_context":{"text":"Tech Updates","link":"https:\/\/areeblog.com\/category\/tech-updates\/"},"img":{"alt_text":"AMD\u2019s Helios AI Rack Marks a New Phase in the AI Infrastructure Race","src":"https:\/\/i0.wp.com\/areeblog.com\/wp-content\/uploads\/2026\/08\/IMG-20260806-WA0025-1.jpg?resize=350%2C200&ssl=1","width":350,"height":200,"srcset":"https:\/\/i0.wp.com\/areeblog.com\/wp-content\/uploads\/2026\/08\/IMG-20260806-WA0025-1.jpg?resize=350%2C200&ssl=1 1x, https:\/\/i0.wp.com\/areeblog.com\/wp-content\/uploads\/2026\/08\/IMG-20260806-WA0025-1.jpg?resize=525%2C300&ssl=1 1.5x, https:\/\/i0.wp.com\/areeblog.com\/wp-content\/uploads\/2026\/08\/IMG-20260806-WA0025-1.jpg?resize=700%2C400&ssl=1 2x, https:\/\/i0.wp.com\/areeblog.com\/wp-content\/uploads\/2026\/08\/IMG-20260806-WA0025-1.jpg?resize=1050%2C600&ssl=1 3x"},"classes":[]},{"id":6488,"url":"https:\/\/areeblog.com\/nvidia-teams-up-with-six-financial-giants-to-mobilize-500-billion-for-ai-infrastructure\/","url_meta":{"origin":6938,"position":2},"title":"Nvidia Teams Up With Six Financial Giants to Mobilize $500 Billion for AI Infrastructure","author":"Daniel Chinonso John","date":"August 11, 2026","format":false,"excerpt":"Nvidia has partnered with six of the world\u2019s largest financial institutions to establish financing platforms aimed at mobilizing more than $500 billion in third-party capital for AI infrastructure, as CEO Jensen Huang argues that AI compute has become an investable infrastructure asset. The partnerships, announced on Monday, August 10, involve\u2026","rel":"","context":"In &quot;Tech Updates&quot;","block_context":{"text":"Tech Updates","link":"https:\/\/areeblog.com\/category\/tech-updates\/"},"img":{"alt_text":"Nvidia Teams Up With Six Financial Giants to Mobilize $500 Billion for AI Infrastructure","src":"https:\/\/i0.wp.com\/areeblog.com\/wp-content\/uploads\/2026\/08\/images-31.jpeg?resize=350%2C200&ssl=1","width":350,"height":200,"srcset":"https:\/\/i0.wp.com\/areeblog.com\/wp-content\/uploads\/2026\/08\/images-31.jpeg?resize=350%2C200&ssl=1 1x, https:\/\/i0.wp.com\/areeblog.com\/wp-content\/uploads\/2026\/08\/images-31.jpeg?resize=525%2C300&ssl=1 1.5x, https:\/\/i0.wp.com\/areeblog.com\/wp-content\/uploads\/2026\/08\/images-31.jpeg?resize=700%2C400&ssl=1 2x"},"classes":[]},{"id":6238,"url":"https:\/\/areeblog.com\/firmus-technologies-signs-nvidia-partnership-to-expand-ai-compute-access\/","url_meta":{"origin":6938,"position":3},"title":"Firmus Technologies Signs Nvidia Partnership to Expand AI Compute Access","author":"Daniel Chinonso John","date":"June 28, 2026","format":false,"excerpt":"According to Reuters on June 28, Australian AI infrastructure company Firmus Technologies has signed a strategic partnership with Nvidia aimed at giving emerging AI firms more affordable access to computing power, according to Reuters. The agreement is designed around Nvidia hardware, Nvidia-powered cloud services, and a revenue-sharing model that ties\u2026","rel":"","context":"In &quot;Tech Updates&quot;","block_context":{"text":"Tech Updates","link":"https:\/\/areeblog.com\/category\/tech-updates\/"},"img":{"alt_text":"Firmus Technologies Signs Nvidia Partnership to Expand AI Compute Access","src":"https:\/\/i0.wp.com\/areeblog.com\/wp-content\/uploads\/2026\/06\/images-3.png?resize=350%2C200&ssl=1","width":350,"height":200,"srcset":"https:\/\/i0.wp.com\/areeblog.com\/wp-content\/uploads\/2026\/06\/images-3.png?resize=350%2C200&ssl=1 1x, https:\/\/i0.wp.com\/areeblog.com\/wp-content\/uploads\/2026\/06\/images-3.png?resize=525%2C300&ssl=1 1.5x"},"classes":[]},{"id":6848,"url":"https:\/\/areeblog.com\/doj-investigates-nvidias-17-billion-groq-deal-over-potential-antitrust-concerns\/","url_meta":{"origin":6938,"position":4},"title":"DOJ Investigates Nvidia\u2019s $17 Billion Groq Deal Over Potential Antitrust Concerns","author":"Daniel Chinonso John","date":"September 10, 2026","format":false,"excerpt":"The U.S. Department of Justice is investigating Nvidia\u2019s deal with artificial intelligence chip startup Groq over concerns that the transaction may have been structured to avoid antitrust review, according to reporting published September 9, 2026. The investigation was first reported by The New York Times, while Reuters reported the development\u2026","rel":"","context":"In &quot;Tech Updates&quot;","block_context":{"text":"Tech Updates","link":"https:\/\/areeblog.com\/category\/tech-updates\/"},"img":{"alt_text":"DOJ Investigates Nvidia\u2019s $17 Billion Groq Deal Over Potential Antitrust Concerns","src":"https:\/\/i0.wp.com\/areeblog.com\/wp-content\/uploads\/2026\/09\/images-56.jpeg?resize=350%2C200&ssl=1","width":350,"height":200,"srcset":"https:\/\/i0.wp.com\/areeblog.com\/wp-content\/uploads\/2026\/09\/images-56.jpeg?resize=350%2C200&ssl=1 1x, https:\/\/i0.wp.com\/areeblog.com\/wp-content\/uploads\/2026\/09\/images-56.jpeg?resize=525%2C300&ssl=1 1.5x, https:\/\/i0.wp.com\/areeblog.com\/wp-content\/uploads\/2026\/09\/images-56.jpeg?resize=700%2C400&ssl=1 2x"},"classes":[]},{"id":6910,"url":"https:\/\/areeblog.com\/alibabas-new-qwen3-8-omni-flash-pushes-ai-agents-into-long-form-video-and-audio\/","url_meta":{"origin":6938,"position":5},"title":"Alibaba\u2019s New Qwen3.8-Omni-Flash Pushes AI Agents Into Long-Form Video and Audio","author":"Daniel Chinonso John","date":"September 20, 2026","format":false,"excerpt":"Alibaba has introduced Qwen3.8-Omni-Flash, a new native omnimodal AI model designed to move artificial intelligence agents beyond understanding audio and video and toward planning tasks, using tools and completing work. The model accepts text, images, audio and video as input and supports a 1-million-token context window. Alibaba says it is\u2026","rel":"","context":"In &quot;Tech Updates&quot;","block_context":{"text":"Tech Updates","link":"https:\/\/areeblog.com\/category\/tech-updates\/"},"img":{"alt_text":"Alibaba\u2019s New Qwen3.8-Omni-Flash Pushes AI Agents Into Long-Form Video and Audio","src":"https:\/\/i0.wp.com\/areeblog.com\/wp-content\/uploads\/2026\/09\/alibaba-qwen_GXn5NKdDXY.jpg?resize=350%2C200&ssl=1","width":350,"height":200,"srcset":"https:\/\/i0.wp.com\/areeblog.com\/wp-content\/uploads\/2026\/09\/alibaba-qwen_GXn5NKdDXY.jpg?resize=350%2C200&ssl=1 1x, https:\/\/i0.wp.com\/areeblog.com\/wp-content\/uploads\/2026\/09\/alibaba-qwen_GXn5NKdDXY.jpg?resize=525%2C300&ssl=1 1.5x, https:\/\/i0.wp.com\/areeblog.com\/wp-content\/uploads\/2026\/09\/alibaba-qwen_GXn5NKdDXY.jpg?resize=700%2C400&ssl=1 2x"},"classes":[]}],"jetpack_featured_media_url":"https:\/\/areeblog.com\/wp-content\/uploads\/2026\/09\/2026-07-16t115806z-1705182164-rc2yemamcyvv-rtrmadp-3-nvidia-japan-2026-07-b91c7076ec607a9fe7bf6bce896062e2.jpg","_links":{"self":[{"href":"https:\/\/areeblog.com\/wp-json\/wp\/v2\/posts\/6938","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/areeblog.com\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/areeblog.com\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/areeblog.com\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/areeblog.com\/wp-json\/wp\/v2\/comments?post=6938"}],"version-history":[{"count":2,"href":"https:\/\/areeblog.com\/wp-json\/wp\/v2\/posts\/6938\/revisions"}],"predecessor-version":[{"id":6941,"href":"https:\/\/areeblog.com\/wp-json\/wp\/v2\/posts\/6938\/revisions\/6941"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/areeblog.com\/wp-json\/wp\/v2\/media\/6939"}],"wp:attachment":[{"href":"https:\/\/areeblog.com\/wp-json\/wp\/v2\/media?parent=6938"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/areeblog.com\/wp-json\/wp\/v2\/categories?post=6938"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/areeblog.com\/wp-json\/wp\/v2\/tags?post=6938"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}