What each tool is actually for, described plainly. No ratings, no scores, no invented benchmarks.
Local models and inference: Running and fine-tuning models on your own hardware, and the servers that make them fast enough to use.
Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)
github.com/hiyouga/LlamaFactoryLocalAI is the open-source AI engine. Run any model - LLMs, vision, voice, image, video - on any hardware. No GPU required.
github.com/mudler/LocalAIGPT4All: Run Local LLMs on Any Device. Open-source and available for commercial use.
github.com/nomic-ai/gpt4allJan is an open source alternative to ChatGPT that runs 100% offline on your computer.
github.com/janhq/janRun GGUF models easily with a KoboldAI UI. One File. Zero Install.
github.com/LostRuins/koboldcppDistribute and run LLMs with a single file.
github.com/mozilla-ai/llamafileGet up and running with Kimi-K2.6, GLM-5.2, MiniMax, DeepSeek, gpt-oss, Qwen, Gemma and other models.
github.com/ollama/ollamaSGLang is a high-performance serving framework for large language models and multimodal models.
github.com/sgl-project/sglangOpen-source desktop app for local LLMs. Text, vision, tool-calling, OpenAI/Anthropic-compatible API. 100% private.
github.com/oobabooga/textgen🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.
github.com/huggingface/transformersLocal UI to run and train LLMs and diffusion models. Supports GGUF, MLX, Qwen3.8, DeepSeek-V4, MiniMax-H3, Gemma 4, FLUX and more.
github.com/unslothai/unslothA high-throughput and memory-efficient inference and serving engine for LLMs
github.com/vllm-project/vllm