How I deployed and optimized local LLMs on a 5-year-old consumer PC using llama.cpp, from hardware bottlenecks to advanced inference techniq...