#thecompaniesapi quantize our new model and bench it on few domains; minimal loss and more than 70% size reduction, we should now be able to fit 2 to 3 instances on our GPU!
#thecompaniesapi properly prioritize user jobs in front of continuous scan; monitor the queue; redeploy textsynth to handle only translation as its much faster; tune ollama config to match system resources better; improve extraction flow; cross check envs and drop old tlds in cloudflare in favor of tailscale in all our infra
#thecompaniesapi test new models on MLX with 8gb/16gb mac minis and it runs!! next up trying q8/q4 to see if it can get to prod; also rewrite lots of our flows parts to support both browser and raw http calls which makes it a lot faster