inari@piefed.zip to LocalLLaMA@sh.itjust.worksEnglish · 23 days agoBenchmarking Qwen3.8 27B quantizations: 4-bit holds up, 1-bit collapsesquesma.comexternal-linkmessage-square11linkfedilinkarrow-up129arrow-down15
arrow-up124arrow-down1external-linkBenchmarking Qwen3.8 27B quantizations: 4-bit holds up, 1-bit collapsesquesma.cominari@piefed.zip to LocalLLaMA@sh.itjust.worksEnglish · 23 days agomessage-square11linkfedilink
minus-squareShimitar@downonthestreet.eulinkfedilinkEnglisharrow-up2·14 days agoGreat model. The best so far for my usage (agents and coding). I can get 15t/s on my dual rtx a4000 setup. I use the Q5, with 300k context.
Great model. The best so far for my usage (agents and coding). I can get 15t/s on my dual rtx a4000 setup. I use the Q5, with 300k context.