LLM inference in C/C++ Written mainly in C++. Released under the MIT licence.
Local models and inference — Running and fine-tuning models on your own hardware, and the servers that make them fast enough to use.
This entry describes llama.cpp in its own words, taken from its own page. KuponGuru publishes no ratings or review scores, and claims no affiliate relationship with any tool listed.