Pinchflat

I Made My Local AI Pick Its Own Model

Raw Attributes

Source: Techno Tim
  • media_id: qxfNTcPl70w
  • media_filepath: /downloads/Techno Tim/2026-10-05 I Made My Local AI Pick Its Own Model/I Made My Local AI Pick Its Own Model.mp4
  • prevent_download: false
  • metadata_filepath: /downloads/Techno Tim/2026-10-05 I Made My Local AI Pick Its Own Model/I Made My Local AI Pick Its Own Model.info.json
  • updated_at: 2026-10-06T22:28:49Z
  • description: I wanted my local AI to choose the right model automatically instead of making me pick between speed and capability for every request. I test NVIDIA Switchyard for local AI model routing with Nemotron 3.5 Lightning on an RTX 3090 and DeepSeek V4 Flash across two GX10s. I cover Stage routing, vLLM, context windows, and some of the problems I ran into while testing it with real coding agents. In my best autonomous coding run, almost 88% of the routing decisions went to Nemotron while DeepSeek was still available when Switchyard decided it was needed. This is pretty close to what I wanted. Most of the work stays on the fast local model, the more capable model gets pulled in when it makes sense, and I don't have to choose between them before I start. Video Notes: https://technotim.com/posts/switchyard-local-ai-routing/ Merch Shop πŸ›οΈ: https://l.technotim.com/shop Support me on Patreon: https://www.patreon.com/technotim Sponsor me on GitHub: https://github.com/sponsors/timothystewart6 Subscribe on Twitch: https://www.twitch.tv/technotim Become a YouTube member: https://www.youtube.com/channel/UCOk-gHyjcWZNj3Br4oxwh0A/join Gear Recommendations: https://l.technotim.com/gear Get Help in Our Discord Community: https://l.technotim.com/discord 2nd channel: https://www.youtube.com/@TechnoTimTinkers (Affiliate links may be included in this description. I may receive a small commission at no cost to you.) 00:00 Automatic AI Model Routing 01:09 Nemotron Lightning vs DeepSeek 01:40 Fixing Nemotron's Performance 02:35 Is Nemotron Actually Good? 02:58 How NVIDIA Switchyard Works 04:55 Putting Stage Routing to the Test 05:28 The 32K Context Problem 06:46 Fixing Routing with 128K Context 07:44 Testing on Real-World Code 08:40 Building the Autonomous Agent Test 09:46 Run 4: Getting Close 10:39 Run 5: 88% Routed to Nemotron 11:50 Run 6: When the Agent Got Stuck 12:11 Is NVIDIA Switchyard Worth It? 13:40 What's Next for My Local AI Thank you for watching!
  • livestream: false
  • id: 7938858
  • matching_search_term:
  • upload_date_index: 99
  • last_error:
  • media_size_bytes: 76797891
  • duration_seconds: 869
  • uploaded_at: 2026-10-05T17:21:32Z
  • playlist_index: 1
  • source_id: 4
  • uuid: 69bf8af4-b233-4761-a351-602139d6c174
  • media_downloaded_at: 2026-10-06T22:28:46Z
  • predicted_media_filepath: /downloads/Techno Tim/2026-10-05 I Made My Local AI Pick Its Own Model/I Made My Local AI Pick Its Own Model.mp4
  • inserted_at: 2026-10-06T03:30:41Z
  • media_redownloaded_at:
  • culled_at:
  • title: I Made My Local AI Pick Its Own Model
  • prevent_culling: false
  • subtitle_filepaths:
  • short_form_content: false
  • thumbnail_filepath: /downloads/Techno Tim/2026-10-05 I Made My Local AI Pick Its Own Model/I Made My Local AI Pick Its Own Model-thumb.jpg
  • nfo_filepath:
  • original_url: https://www.youtube.com/watch?v=qxfNTcPl70w
Worker
State
Scheduled At
Pinchflat.Downloading.MediaDownloadWorker completed