HN in RSCserver-reason-react
top.mdnew.mdbest.mdask.mdshow.mdjobs.md
← Back to stories

Strands Decider 2B: a small, open-source, decision model

186 pointsby gmays 8 hours ago46 comments

Discussion

Loading discussion
  • miguelspizza · 8 hours ago

    This is a great model. I've been running it on device in chrome extension to filter things like email. It is just the right mix of size, capability and speed to make it generally useful for adhoc bulk classification tasks. For those wanting to run it in browser: https://huggingface.co/alxnahas/strands-decider-2B-webgpu

    • babelfish · 7 hours ago

      Live demo: https://alxnahas.github.io/strands-decider-web/?backend=engi...

  • keyle · 8 hours ago

    Fantastically well written. It's rare for me to be able to understand what the AI gurus are talking about, and this was written by humans for humans. It can technically be used for a lot of use cases, I'd like people to chime in on ideas on this?

    • nryoo · 6 hours ago

      Picking lunch menu..?

    • mattstir · 2 hours ago

      A use case I have in mind is creating a relatively simple home assistant app that would let me control smart home things like lightbulbs and playing music on some kitchen speakers. I can also imagine having it differentiate between commands like "set lights to 30% brightness" and more general questions like "what's the weather today" and piping the latter over to a cheap LLM to handle. It could probably all be handled by an LLM, but I like how fast these decision models end up being.

  • adenta · 7 hours ago

    At this point I can't wait for a comedian to release a decision model backed by humans. Meet Jerry- it's literally a guy named Jerry answering your questions.

    • melvinram · 6 hours ago

      Your comment made me chuckle. Jerry, "The Decider" https://www.youtube.com/watch?v=r8VbzrZ9yHQ

      • bfung · 6 hours ago

        Same, but my mind went straight to Rick and Morty… def don’t want Jerry deciding XD

        • throwawayffffas · 1 hour ago

          Ah, easy to work around, you just iterate eliminating his choices one at a time, until you get the one he didn't pick (that's your correct pick).

    • Balooga · 6 hours ago

      Oh! You probably want ChatTJB [1] [1] - https://chattjb.org/about

    • jeef_berky · 5 hours ago

      Don't worry, Jerry will rig everything up for you.

    • jeremycarter · 5 hours ago

      RACE HIM JERRY!

  • davvie · 6 hours ago

    Looks really nice, I think I could use it on my Mac mini for some smaller automations

  • teruakohatu · 6 hours ago

    Any idea how well this would run on a CPU?

    • gopalv · 5 hours ago

      On my M3 mac, it works okay inside a docker container with just CPU. { "model": "strands-decider-2B-hobson-v19", "answers": { "is_urgent": { "type": "noul", "noul": 0.8287 } }, "usage": { "input_tokens": 86, "output_tokens": 1 }, "latency_ms": 1732.17 } This is how I got it running - https://gist.github.com/2891eb0db9ea92c1a4e860d44f556292 There's a lot more to be done if we optimize for MLX & let it run on a Mac mini instead of the docker wrapper.

    • avereveard · 5 hours ago

      About half a second per decision on six cores

      • davidwritesbugs · 3 hours ago

        Isnt that a bit slow for these?

  • soltanov · 6 hours ago

    Benchmark calibration does not establish reliability on unfamiliar production inputs.

  • yieldcrv · 5 hours ago

    a strand type game

  • SubiculumCode · 5 hours ago

    Are any of these multimodal yet? I'd love to try asking a model with calibrated probabilities to answer question like, "do these shapes match?". Sure, you can ask a LLM....

    • necubi · 5 hours ago

      Cloudflare’s clef is multimodal ( https://blog.cloudflare.com/clef-decision-models/ ) (Disclaimer, I work at Cloudflare, but not on models)

    • sauhsoj · 5 hours ago

      Strands Decider can take vision in. How does it go with that question?

      • Zopieux · 2 hours ago

        This is not advertised on their page, did you make this up? I believe image classification/analysis by deciders (not just OCR, not everything is about text) is still lacking. Cloudflare's Clef had fair results on my test, but it's larger and slower. Wondering about Strands.

  • stephantul · 5 hours ago

    2B being called small is such a sign of the times

  • mattvr · 5 hours ago

    Why is everyone calling binary choices `noul`? Does this have some meaning or is it just copying Jev’s API?

    • haarts · 5 hours ago

      It's from Bernoulli.

      • jtfrench · 4 hours ago

        Interesting. Is that a unit he invented or is it just a reference to his last name that stuck?

        • mijoharas · 4 hours ago

          From Bernoulli maps apparently. I still don't understand why[0]. [0] https://news.ycombinator.com/item?id=49723267 (see parent for reference)

          • henrymerrilees · 2 hours ago

            A bit of a garden path path sentence... I believe "maps" in "maps to" was being used as a verb. So "noul"--short for Bernoulli as in the Bernoulli distribution--"maps to if-statements."

            • mijoharas · 1 hour ago

              oh god, you're completely right! nice and obvious on a reread. (I also just learnt about "garden path" sentences.) I'd googled "bernoulli map" after reading that, and came across this[0], which I wasn't aware of and thought was somehow related so I entrenched my misunderstanding (I didn't dig in.) Side note: I just wrote the sentence "you're completely right!", and almost changed it because it sounds like slop now. I wonder if we're gonna get an increase in these sentences in human written prose over time as people mimic these sentences, or a decrease as people shy away from them to not sound like AI. :) [0] https://en.wikipedia.org/wiki/Dyadic_transformation

      • fennecfoxy · 1 hour ago

        Huh. And here I was thinking the obvious thing is that it's a way to have "null" without it accidentally being parsed as null.

    • WASDx · 3 hours ago

      The new OpenAI Decisions API calls it "predicate". Also calling the API "decisions" rather than "system one". Usually I don't like inventing new standards but I hope the OpenAI schema takes over. We don't need this hype terminology.

    • dprkh · 3 hours ago

      Claude generated it and it stuck.

    • tchalla · 3 hours ago

      The same reason why they’re calling this System One thinking. Everyone wants to be seen doing different things and smart ones.

    • hiimkeks · 2 hours ago

      My guess is if you pronouce this "nool" it sounds similar to "bool", and it's a slice of the name "Bernoulli" because the models generate Bernoulli distributions

  • hrpnk · 4 hours ago

    clef from cloudflare runs on llama.cpp - being locked-in to strands cli would be a bummer and will slow down adoption. Since it's a LoRa on Qwen, I assume this is runnable via llama.cpp. Pity that the PEFT/LoRa->GGUF translation is left to the user. Anyone got past: $ uv run --with transformers==5.19.0 convert_lora_to_gguf.py ~/Downloads/lora --dry-run --verbose [...] File "/Users/user/repos/llama.cpp/conversion/base.py", line 630, in map_tensor_name raise ValueError(f"Can not map tensor {name!r}") ValueError: Can not map tensor 'layers.0.linear_attn.in_proj_a.weight'

  • real_faxenoff · 3 hours ago

    As a regular user of a bunch of specialized micromodels, I'll tell you this: you won't be happy with such a model (and its JEV counterparts) running permanently in the background on your PC's CPU. You need to offload their processing to the NPU. There are many pitfalls along the way, but the result is worth it. NPU performance will be twice as high, while power consumption will be four times lower. No additional fan noise (if you know what I mean). I'll wait another month until the first phase of the =battle royale= among models of this kind wraps up, put together a solution for the NPU/iGPU, and post it on HF.

    • cyanydeez · 8 minutes ago

      I haven't seen any frameworks for running the NPU. my 395+ needs a buddy.

  • woadwarrior01 · 3 hours ago

    The ~2-week-old Intern-Decision family of models (0.8B, 2B and 4B) have the same Qwen3.5 base model family (albeit the instruction-tuned variants) and pointer head architecture. https://huggingface.co/collections/internlm/intern-decision

  • mynti · 3 hours ago

    Can someone explain this architecture a bit more in depth? They say the pointer head scores the hidden state at each option against the hidden state of the answer. But the LLM produces hidden states per token, so an option can span multiple tokens, no?

  • girvo · 3 hours ago

    Does anyone know if it is worth fine-tuning one of these decision models on the shape of the questions you want it to work on, vs the more general versions? I'm using Jev pretty successfully at work at the moment, but am curious about what is doable

    • weinzierl · 2 hours ago

      I'm interested in this as well and maybe to broaden the scope of the question a little: If I have a sizeable amount of labeled data and need decisions calibrated to that data should I 1. Ignore the hype and train a traditional classifier 2. Finetune an LLM based decision model 3. Shoehorn (probably a small subset of) the data into the context of the LLM classifier somehow If the answer is 3. where does the data belong? In the input content? Request wide state? In the question instructions? In the criteria? How much of my data can and should I use?

    • gxcsoccer · 2 hours ago

      I think jev points to an interesting way to fine-tune open models The goal isn’t to replace current models, it’s to train a model for a specific domain so it can handle multiple-choice and yes/no questions quickly, helping the overall system run faster and get better results

  • Ujj-001 · 2 hours ago

    is jev commoditized now ?

  • kimseungyong · 2 hours ago

    I wait for this open-source model. I should let the development agent select a model and run it to improve token efficiency. Is it being used this much these days?

  • ContinuityLab · 1 hour ago

    Swapping out the text-generation head for a dedicated pointer head on a small footprint model is a pragmatic approach for low-latency local decision pipelines.