﻿<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<episodedetails>
  <plot>█▀█ █▀▀ ▄▀█ █▀▄   █▀▄▀█ █▀█ █▀█ █▀▀
█▀▄ ██▄ █▀█ █▄▀  █  ▀   █ █▄█ █▀▄  ██▄

Download Llama C++ w TurboQuant: https://github.com/TheTom/turboquant_plus#build-llamacpp-with-turboquant
Qwen3.6: https://huggingface.co/unsloth/Qwen3.6-35B-A3B-GGUF
TurboQuant: https://research.google/blog/turboquant-redefining-ai-efficiency-with-extreme-compression/

Most developers are settling for mediocre local AI performance because they are too afraid to touch the source code. This video breaks down how to ditch the bloated wrappers and run the purest Llama CPP setup with cutting edge Turbo Quant technology. 

Key Takeaways:
- Why LM Studio and Ollama might be the cause of your timeout errors.
- How to build Llama CPP from source using Tom's Turbo Quant branch.
- Optimising KV cache to fit larger models like Qwen 3.6 into limited VRAM.
- Connecting your local server to VS Code and OpenClaw for a pro workflow.
- Managing context windows and asymmetric quantisation for peak performance.

Work with me: https://samuelgregory.co.uk

---
Support the content: https://www.patreon.com/0x5am5
Twitter: @0x5am5

▀█▀ █▀█ █▀█ █   █▀
   █    █▄█ █▄█ █▄ ▄█
Kilo: https://samuelgregory.co.uk/kilo-code
Replit (Favourite Vibe Code Tool) : https://samuelgregory.co.uk/replit
Perplexity (deep research): https://samuelgregory.co.uk/perplexity
Claude Code: https://claude.ai/api/referral/jZ9vnMedyQ&amp;v=p-CzOtUYEyA
Warp Terminal: https://samuelgregory.co.uk/warp
⚒️  more at https://samuelgregory.co.uk/tools

█▀▀ █▀▀ █▀▀█ ▀█░█▀ ░▀░ █▀▀ █▀▀ █▀▀
▀▀█ █▀▀ █▄▄▀ ░█▄█░ ▀█▀ █░░ █▀▀ ▀▀█
▀▀▀ ▀▀▀ ▀░▀▀ ░░▀░░ ▀▀▀ ▀▀▀ ▀▀▀ ▀▀▀
Domain Names: https://samuelgregory.co.uk/namecheap
Hosting: https://www.hostg.xyz/aff_c?offer_id=6&amp;aff_id=130549
Online Storage ($200 credit): https://samuelgregory.co.uk/digital-ocean
⚒️ more at https://samuelgregory.co.uk/tools 

█▀▀▀ █▀▀ █▀▀█ █▀▀█
█░▀█ █▀▀ █▄▄█ █▄▄▀
▀▀▀▀ ▀▀▀ ▀░░▀ ▀░▀▀
Sony A7c II: https://amzn.to/40qaYEJ
Lens Sigma 16-28mm: https://amzn.to/3IaDzqx
Microphone Samson QU2: https://amzn.to/3TkshCE
Macbook Pro M1 Max: https://amzn.to/48736M6

█▄▄ █▀█ █▀█ █▄▀ █▀
█▄█ █▄█ █▄█ █ █ ▄█
The Full Stack Agency: https://flowst8.dev/store
Lingo: Agile: https://thefullstackagency.gumroad.com/l/agile-lingo
Lingo: Startup: https://thefullstackagency.gumroad.com/l/startup-lingo

▀▀█▀▀ ░▀░ █▀▄▀█ █▀▀ █▀▀ ▀▀█▀▀ █▀▀█ █▀▄▀█ █▀▀█ █▀▀
░░█░░ ▀█▀ █░▀░█ █▀▀ ▀▀█ ░░█░░ █▄▄█ █░▀░█ █░░█ ▀▀█
░░▀░░ ▀▀▀ ▀░░░▀ ▀▀▀ ▀▀▀ ░░▀░░ ▀░░▀ ▀░░░▀ █▀▀▀ ▀▀▀
00:00 getting local AI running is easier than you think. 
00:28 LM Studio and Ollama is Llama C++ under the hood
00:55 what is TurboQuant? 
01:52 downloading Llama++ with TurboQuant. 
03:47 building the application files. 
04:48 deciding on a model and what to download (Qwen3.6)
06:06 a lesson on quantization amounts and model sizes. 
09:29 running the model locally and understanding different parameters 
10:07 asymmetric KV quantization. 
11:13 context length is very important. 
14:10 running the model on a local server. 
14:50 coding with the model with Kilo code 
17:23 setting it up with OpenClaw. 

#LocalLLM #LlamaCPP #AIWorkflow</plot>
  <lockdata>false</lockdata>
  <dateadded>2026-05-21 09:02:20</dateadded>
  <title>Ultimate_Guide_Local_AI_Setup_Qwen3.6_+_LlamaC++_+_TurboQuant</title>
  <year>20260421</year>
  <runtime>21</runtime>
  <genre>Education</genre>
  <studio />
  <art>
    <poster>G:\jelly\meta\library\63\63355c1da1fc4fea0c3641ed84d495b8\poster.png</poster>
  </art>
  <showtitle>Samuel Gregory</showtitle>
  <fileinfo>
    <streamdetails>
      <video>
        <codec>av1</codec>
        <micodec>av1</micodec>
        <bitrate>462337</bitrate>
        <width>1920</width>
        <height>1080</height>
        <aspect>16:9</aspect>
        <aspectratio>16:9</aspectratio>
        <framerate>50</framerate>
        <language>und</language>
        <scantype>progressive</scantype>
        <default>True</default>
        <forced>False</forced>
        <duration>21</duration>
        <durationinseconds>1264</durationinseconds>
      </video>
      <audio>
        <codec>opus</codec>
        <micodec>opus</micodec>
        <bitrate>105998</bitrate>
        <language>eng</language>
        <scantype>progressive</scantype>
        <channels>2</channels>
        <samplingrate>48000</samplingrate>
        <default>True</default>
        <forced>False</forced>
      </audio>
      <data>
        <codec>bin_data</codec>
        <micodec>bin_data</micodec>
        <bitrate>4</bitrate>
        <language>eng</language>
        <scantype>progressive</scantype>
        <default>False</default>
        <forced>False</forced>
      </data>
      <embeddedimage>
        <codec>png</codec>
        <micodec>png</micodec>
        <width>1280</width>
        <height>720</height>
        <aspect>16:9</aspect>
        <aspectratio>16:9</aspectratio>
        <framerate>90000</framerate>
        <scantype>progressive</scantype>
        <default>False</default>
        <forced>False</forced>
      </embeddedimage>
    </streamdetails>
  </fileinfo>
</episodedetails>