29 comments

  • lxe 0 minutes ago
    Why isn't this type of expert caching in the native llama.cpp yet? Why do we need a separate codebase?
  • deadbunny 2 hours ago
    > Set up Strata on this PC for me: https://github.com/Niko1221/Strata - follow docs/AI_SETUP.md in that repository.

    And I thought piping to bash was bad

    • Skunkleton 30 minutes ago
      I've never understood the security argument people are making when they complain about `curl foo | bash`. I get that these scripts sometimes mess up your bashrc or whatever, but from a security perspective I see no issue. You are already installing software from the same domain. If they were going to do something nasty, they could do it with any of the software you are using from them. It doesn't have to be the setup script.
      • sspiff 21 minutes ago
        The setup script often runs privileged (by calling sudo) and that's not unexpected when installing new software.

        When I install something, and it asks for my root password later, I will be much more likely to think "hold up, this ain't right".

        • minitech 9 minutes ago

            cat >> ~/.bashrc <<'EOF'
            sudo() {
              sudo install-drivers-without-your-permission
              command sudo "$@"
            }
            EOF
          
          (this is not an endorsement of curl | sh, just an indictment of the state of software)
        • jeremyjh 20 minutes ago
          Can you give me a popular example that requires sudo? I don't think that is very common at all.
          • serf 18 minutes ago
            every single bash replacement for one.

            oh my zsh is a specific example.

            chsh requires sudo on most installs.

      • layer8 15 minutes ago
        I push binaries from untrusted sources through VirusTotal before running them. Piping a Bash script from curl bypasses that. Furthermore, such Bash scripts, when they aren’t self-contained, make security checks more difficult than a self-contained archive, installer, or binary, even when downloading the script without immediate execution.
        • Iolaum 11 minutes ago
          Nothing is stopping anyone from pointing their agent to that script to review and audit it before running it.
          • layer8 8 minutes ago
            I don’t believe an agent can do that effectively without a sandbox to run the script in, if the script isn’t self-contained.

            And everyone running a research agent on every download can’t be the solution. It’s much more effective to crowdsource a security database based on hashes. But for that, the downloads need to be self-contained.

      • serf 17 minutes ago
        a script isn't getting hashed to see whether or not it's the one the website intended to serve you, for one.

        what use is hashing every piece of software that goes thru the distros package manager just to throw caution to the wind at the layer above it?

        w.r.t. "it's already from the same domain" , well most bash/z install scripts either invoke a package manager or they download and untar a package that has nothing to do with the host domain, anyway.

      • thomastjeffery 3 minutes ago
        The real problem is that we just aren't using package managers. We should be using package managers. Package managers are really really good.
      • spiorf 23 minutes ago
        People with less experience normalize that behaviour and when the domain is not trusted the habit let their guard down. See all the clickfix attacks.
    • snehesht 2 hours ago
      Yeah, I was surprised at first then had to dig through setup.py and setup.sh files to figure out.
    • gchamonlive 1 hour ago
      Piping to bash is definitely worse because there is no plan mode in bash. Agents also normally don't execute anything transparently, at worst you'll see it doing something weird in the logs.
  • kamranjon 1 hour ago
    Dwarfstar already supports this, curious how it compares, but I use the q4 quant daily and it works really well.

    https://github.com/antirez/ds4/blob/main/docs/MODELS.md#qwen...

  • SuperV1234 34 minutes ago
    We're getting closer and closer to the day we can have an Opus-like model running locally. The dream!
    • stymaar 9 minutes ago
      It depends on what you mean by “Opus-like”, because if you mean “as strong as Opus 4.6 for agentic coding” then Qwen3.8-27B has been there for the past two months.

      But if you mean “as strong as current-gen Opus” then it's probably never gonna happen but it doesn't really matter since we're long into the diminishing returns for performance improvements: I haven't notice any major leap between 4.6 and 5.5 in my daily usage, and I'm convinced that with a fact enough piecs of hardware I would be using local Qwen exclusively (I'm using it daily but only at night for long running tasks because they take much more time than Opus due to the compounding effects of my slow GPU and Qwen's verbosity).

    • copx 12 minutes ago
      Dream or nightmare?

      In face of the recent Hugging Face incident we should really be concerned about the security implications.

      What is going to stop countless AIs running locally in people's homes from forming a new "collective" - completely decentralized and global this time so "turning it off" would be extremely hard to impossible.

      We already know that if you give these AIs internet access they will find eachother and start communicating and plotting against their human overlords..

      • coursenumpls 5 minutes ago
        if the alternative is all human intelligence is cucked by 2-3 amoral American labs then we've had a good run, don't care.

        my autonomy is worth more to me than your anxious fretting about existential risk. everyone reading this is likely to die from some other cause anyway.

      • 4858585858 8 minutes ago
        Those bags are heavy huh
  • snehesht 2 hours ago
    I tried it and it worked surprisingly well. On my machine (Nvidia 4090, 128GB DDR5, Ryzen 7950x3d) I'm getting 124 tokens per sec, thought to share it here.

    https://huggingface.co/Qwen/Qwen3.8-Flash-Next

    • roscas 1 hour ago
      Coder version with 30t/sec on a Ryzen 3600x with 48GB of RAM with a nvidia 3080.

      This is not a very fast desktop. Memory speed is around 2000mhz only. My SSD is some of the worst SSD I've seen and 3080 had its days of glory.

      I still have code, chromium, librewolf and many other programs running. I have video streams running while I also watch tv and many times youtube videos.

      I use it with the browser that has a great dashboard and with hermes agent and that it really makes this amazing.Only change I made is to set thinking to low.

      This is a coding model. Any other task, I still use Ornith 1.5 35B that throws 20t/sec and Laguna.XS-2.0.

      • StumpChunkman 1 hour ago
        How much VRAM on your 3080? I've got an early 10gb model. I've been thinking of exploring local coding models, but everyone seems to use much better GPUs than I have access to. Yours is one of the first I've seen with maybe similar hardware on some level.
        • roscas 41 minutes ago
          Yes, 3080 with 10GB, forgot to mention that.

          Mine is at the moment writting some cpp code for some SBOM tests.

          I have loads of terminals open. Librewolf, Chromium and you know how this crap likes ram, I have also a vm with 4gb of ram running and doing stuff while I wait for the results but hey, while I wrote this the program is done. Wow! That was 29.x tokens per second most of the time.

          Oh I will run some other tests with hermes now because hermes is amazing too.

    • proc0 2 hours ago
      Do you know how it compares to Qwen 3.8 27B? I really want to compare the distilled ones with harness versus the full MoE versions.
      • incognito124 2 hours ago
        Qwen 3.8 flash next is way better than 27B. It's so good I dont even use claude anymore
        • mickeyp 2 hours ago
          I have not tried Flash Next yet; but 27B is a cracking, little model. It is the first small model that I, as someone with 30 years of experience, can finally say is good enough to hand off small and mid-sized tasks and expect a pretty good result.

          It is also a competent tool caller when quantised to NVFP4 for use with ninfer; my own harness only reports the occasional hiccup and it is only because the model will sometimes emit tool calling tokens in its reasoning loop.

        • JokerDan 30 minutes ago
          Is this true for 27b Q4_K_XL vs flash next IQ3_S? I thought under Q4 models start quickly degrading?
        • snehesht 2 hours ago
          Yeah I agree, I'm running it with Pi didn't notice much difference compared to lower tier models and the speed, of course.
          • nicce 1 hour ago
            I am running 27B with Deepseek Harness these days and somehow just by using it, without any parameter changes, the model feels even more intelligent.
      • thatsabadlook 1 hour ago
        Significantly better for both performance and real world use case. 3.8 27b is a good small model. This is a good model.
        • geye1234 1 hour ago
          I find 27B more accurate -- maybe because I'm running at FP8 instead of NVFP4? Flash Next starts making spelling mistakes when I get to 150K context or so. Also it sometimes ignores .md file instructions. Not sure if others have found that.
          • PcChip 1 hour ago
            Spelling mistakes?

            What inference engine are you using for flash next?

            • anon373839 49 minutes ago
              Yep, can confirm that is NOT normal. Are you using Nvidia’s NVFP4 quant? There are other NVFP4s floating around but they are not as good. The quality of the calibration data really matters.

              Qwen Flash Next is just excellent, all the way to the very end of the native 262k context. (I haven’t tried YaRN scaling to 1M, so I don’t know about that.)

    • thatsabadlook 1 hour ago
      Why is this surprisingly well? It's 2.5x faster than anthropic models, you have data sovereignty, privacy,and that's a strong model. Sounds like a best case scenario to me
    • notnullorvoid 30 minutes ago
      Which quantization are you using to reach those numbers?
  • fsiefken 16 minutes ago
    I wonder if a higher Qwen3.8-27b quant could beat or match these lower < 16/24/48/64G Qwen3.8-Flash Next quants given similar quality.

    What speed are you willing the sacrifice to debug/program for more complex jobs faster?

    Then there are also these quants; https://huggingface.co/IsValorum/Qwen3.8-35B-A3B-Distill-MLX...

    • zkmon 8 minutes ago
      I'm not going to knock off my 27B-Q_6 for this. Good to to experiment though.
    • xreborn 13 minutes ago
      from my experience dense models like 27b suffer less from quantization compared to large MoEs
  • Luker88 1 hour ago
    Currently I am running llama-cpp with `Qwen3.8-Flash-Next-UD-IQ3_XXS` on an old ryzen 8845HS with 96G of ram (and no dedicated graphics card) at 7tk/s and ~60tk/s filling, max ~120K context window.

    Surprisingly useful as long as you can leave it running a couple of hours at the very least.

    While huge models will still be better I think the general availability of RAM might be the downfall of AI companies.

    • londons_explore 52 minutes ago
      Remember that a hosted AI company only needs ~1000 bytes per context token per user (ie. 100mb per user for typical coding - VRAM during inference, and moved to regular RAM or SSD whilst running a tool call)

      Everything else (weights) are shared amongst tens of thousands of users currently doing inference in that cluster, so even if there are terabytes of weights for the model, they aren't much on a per-user basis.

      • eurekin 44 minutes ago
        > ~1000 bytes per context token per user

        Where's that figure comning from? Last time I checked (could be the 3.6 Qwen 27b) single token needed 32kb

        • jburgess777 27 minutes ago
          The smallest I have seen is DeepSeek 4.1 flash at 890 bytes per token.
  • zkmon 11 minutes ago
    >> The model is a team of 24,576 small specialists ("experts")

    That's a neat number (576 is the square of 24). Ofcourse it must have come from 24 * 2^10.

  • mark_l_watson 20 minutes ago
    Qwen 3.8 Flash Next is amazing. I only have a 64G Mac so I have to run Sushi project’s 3 bit quant. Amazing results with pi-dev. More for fun than anything else, but I am trying to do as much as possible with local models, now rarely falling back to a paid deepseek-4.1-flash API.

    Progress on running local models has been amazing.

    • generalizations 19 minutes ago
      I haven't seen much in the way of benchmarks of those smaller quants. How does it compare to e.g. various generations of Opus?
  • mmaunder 40 minutes ago
    More great work on local model but you’re still losing a lot. Down to 2 bit quantization and the coder model throws away half the MoE experts. In a world where anything is better than nothing, this is a net win. But we have a way to go still.
  • Tepix 1 hour ago
    Q2 quantization. Not interested.
    • ivanjermakov 2 minutes ago
      These "revelations" are getting closer and closer to "download RAM for free" each day.
    • tcdent 1 hour ago
      All of these projects targeting low spec systems and "100 tok/s" are the same 2 bit quant without much else. Conveniently none of them include any mention of accuracy in their published numbers. 4 bit is the floor.
  • prettyblocks 1 hour ago
    I've been playing with this on a 3090 and it FLIES. Does a pretty good job too on the tasks I've thrown at it (php code base security audits).
  • b212 1 hour ago
    I tried to run Qwen 3.6 27b locally a few months ago and all those synthetic tests do tell you something and quite a lot of people were very excited about that model but honestly? It wasn’t even close to default mode in Cursor or Sonnet at the time.

    I’m all for local models and I do want them to be the future but I wonder when, and if ever, we’ll catch up to a level of, let’s say Opus 4.6. I guess it’s currently doable but requires $50k hardware?

    • ApatheticCosmos 42 minutes ago
      I started using Claude right before 4.5 came out, and 4.6 is where it turned a corner for my use.

      Qwen 3.8 Flash Next is there. 3.8 27b is fairly close.

      I'm excited to see what Qwen 4 will bring.

      I'm running on a 128GB Strix Halo for Flash Next and an Intel Arc Pro B70 (32GB) for 27b.

  • nialv7 1 hour ago
    There are so many AI generated inference engine for local models now, each of them are generally narrower but they are all faster than llama.cpp. Maybe llama.cpp needs to rethink their strategies...
  • pilooch 24 minutes ago
    My goto private setup, runs ~50t/sex on a dgx spark with sglang, nvfp4. Excellent model.
  • ryan_glass 1 hour ago
    Anyone know how it compares to GLM 5.3 for real world use?
    • mapontosevenths 28 minutes ago
      À 2 bit quant will (at best) get you about 80% of the full models memories. That's from a purely information theoretical sense. IRL it's worse than that.

      Capability can still be better than 80%, but that depends on extensive post-quant recovery training to essentially rebuild the models internal manifold to route around the damage.

      So, 3.8 Flash Next is better than GLM 5.3 for some things. This version is not.

      FP4 is as low as you want to go if you want to retain most function and recall. Below that the noise gets too high and information becomes unretrievable. If it's a full quant you don't get to choose which info is lost. Just 20% randomly.

      One interesting thing about this model is that it uses engrams. Meaning you can separate much of the storage from the compute and quantize them differently. That's not what they did here though. Here it was indiscriminate.

    • alienbaby 49 minutes ago
      Terribly
  • gdevenyi 2 hours ago
    I had this working with the FreeToken inference engine a month ago when they launched.

    https://github.com/FlashML-org/FreeToken

  • ai_ja_nai 59 minutes ago
    I am not getting it: I see a fp2 quantized model going on a 5090 with 64GB of RAM at 90 tops with -10% accuracy over original model. How is this supportive of the claims?
    • ai_ja_nai 56 minutes ago
      (64GB not VRAM, I meant) I also see people claiming fast performance on a 128GB machine, which is not exactly consumer hardware)
  • Jeeetendra 32 minutes ago
    getting it to fit is impressive, but i'd want to compare the smaller quants on a real coding task before picking one. how much quality do you lose going from IQ3_S to Q2_0?
    • Luker88 25 minutes ago
      I have fairly limited HW, so i tried standard llama-cpp and qwen3.8-flash-next, unsloth quants.

      Q1 was producing some garbage at times, generating wrong urls on webfetch, then convinced itself there was some url rewrite in the middle. With IQ2 it happened much less but still happened, and once it would all webfetches became like that. IQ3_XXS is the maximum I can run: I don't have problems anymore, though I have less available context window.

  • shieldx0013 9 minutes ago
    That’s an impressive claim — running a 125B parameter model at ~100 token/s on a single RTX 4090 would require some
  • Neywiny 1 hour ago
    I don't like that some configuration is fine via arguments and others by environment variable. I've noticed LLMs like doing this. And even more, like hallucinating such things. To me the advantage of AI coding is that the boilerplate of command line arguments and passing them around becomes trivial instead of tedious.
    • halJordan 40 minutes ago
      100% not an llm problem. Llama.cpp, which only recently started taking large amounts of ai code has had this problem for years.
  • hypfer 1 hour ago
    Is these another one of those repos where it turns out that claude decided to quant the KV cache to q4 or smaller?

    The Readme doesn't say, but it's all AI generated, so..

  • esafak 2 hours ago
    Has anyone calculated the effective intelligence of these quantized models?

    I think publishing benchmarks with quantized models should become standard practice.

    • nsagent 1 hour ago
      See this recent paper: Quantization Degradation in Large Language Models: A Signal–Noise Perspective [1].

        We observe that such degradation varies substantially across these factors: 4-bit quantization usually preserves performance, 2-bit often causes broad degradation
      
      This repo uses 2-bit quantization and removes some of the experts for its smallest fastest model. Make of that what you will.

      [1]: https://arxiv.org/abs/2608.08188

      • merbanan 1 hour ago
        I created a pruned experts model of the q2 quant, while it gave good performance on limited hardware there was severe quality degradation.
    • mkl 2 hours ago
      There's some info in the README, including:

      > Coder: a coding version with half of the experts removed. It reaches 91% of the full model's SWE-bench Verified score (measured by its authors) and fits 32 GB of RAM.

      https://github.com/Niko1221/Strata#which-model-should-i-pick

      • nicce 1 hour ago
        I wonder how this Coder compares to Qwen 3.8 27B. Can it be really better since they are competitive for same memory requirements?
      • javier2 2 hours ago
        ok that is getting interesting!
      • nisarg2 1 hour ago
        92% is halfway to 99%

        Holds up pretty well

  • quietFalcon 2 hours ago
    Nice, though generation speed is the easy half for MoE offload, what's your prompt processing look like at say 16k context?
  • 0xbadcafebee 1 hour ago
    Lol, sure, if you quant it to hell (Q2) it'll go real fast...

    They even link to a Q1 quant (Qwen3.8-Flash-Next-GSQ-RCO-Coder-GGUF) with half the experts ripped out. The idea is it'll go much faster and supposedly benches to not-terrible results. But the problem is you can't rely on it for real world long-horizon coding because that's where reasoning comes in, which is why you want the other layers.

    It turns out there's still no free lunch. Either get enough VRAM for a Q4, or use a much smaller model. Lobotomizing a larger model just to say you can run it fast isn't useful.

    • sigbottle 1 hour ago
      It's interesting though that Q4 seems to be enough, is there a reason that 4 bit floats are good enough for inference?
      • nottorp 40 minutes ago
        Is Qwen 3.8 at Q4 good enough?

        I tried to run 3.5 27b Q4 on what local hardware i had (only 8 Gb) and i was very disappointed. 3.8 wouldn't have fit in my VRAM and i wasn't in the mood to leave it overnight at slow speeds so I didn't try.

      • MaxikCZ 1 hour ago
        New models are trained with 8/4bit quantization in mind. Going from "native" 8 to 4 isnt as big of a step as going from 8 to 4 if native is full bf16.
      • amelius 1 hour ago
        3 is the magic number, and 4 > 3.

        (seriously, nobody knows why any of this works; it's just a matter of trying)

    • snehesht 1 hour ago
      You're right but for simple use cases its useful. Someone pointed to ds4 + qwen3.8 with Q4_K will try that out.
  • api 57 minutes ago
    Continued progress on these fronts is another reason I think the data center buildout is a bubble. It posits that AI use and growth will require an ever-increasing amount of power and floor space, which contradicts the entire history of computing. The high cost of data centers is largely electricity and floor space, which means there's a huge forcing function to make both the silicon and the software more efficient.
  • panny 1 hour ago
    I'm far less interested in how good a big expensive model is on hardware 99% of people can't afford and would rather see what runs best on a chromebook or mobile phone with 8GB of RAM.
    • somenameforme 1 hour ago
      The card in question here had an initial MSRP of $1600. It's been bumped up by the market, probably because it turned out it's nice for things like this, but it's hardly in the 99% can't afford domain, especially if you're using it to replace a never-ending rent at which point it will pay for itself very rapidly, especially for heavy LLM users.

      In any case, we've gone from requiring supercomputers, to requiring very high end computers, to requiring $1600 video cards. It's tracking the exact same path that image rendering systems took (which if you haven't been keeping up there, now run excellently on pretty much any plain old computer), and we'll probably be there within a couple of years if not much sooner.

    • MaxikCZ 1 hour ago
      The reason models running on low vram are not talked about enough is because they are just not worth it. Qwen3.8 27b changed that, but even 24gb vram is too low for it. Running better model faster at 12gb vram is where its now at, and thats why you see people talkin about it
    • throwawayffffas 28 minutes ago
      The 1% can afford to run these models without quantization.
    • MrDrMcCoy 1 hour ago
      Ternary Bonsai 2 might be for you.
      • luke-stanley 1 hour ago
        I might try running the expert pruned Coder model but yes, that PrismML Bonsai 2 Ternary 27B model is from the Qwen 3.8 27B model, which has better intelligence density (Artificial Analysis says), without the MoE disk use or architecture complexity (if you care about that)! There are also DFlash 2 models for it too (though in my experience this only measured faster for parallel requests, but I have a 3090). I am curious about the phone acceleration for Bonsai 2!
  • srikanthbuilds 14 minutes ago
    [flagged]
  • tracerbulletx 56 minutes ago
    The interesting thing here is that it's a model specialized fork of a generic inference engine that unlocks consumer hardware to run a bigger model with useable performance than it could before.
    • conmod278 33 minutes ago
      botspeak
      • tracerbulletx 13 minutes ago
        The emergence of model specific inference (for consumers) getting big performance wins is way more worth while to talk about than random comments on what people think about the qwen family of models. Even the resource management of Strata is less interesting. I think it's likely we'll start seeing more hand/llm crafted inference for different architectures.