llama_cpp_for_radxa_dragon_wing_q6a

History

Georgi Gerganov 196f5083ef common : more accurate sampling timing (#17382 ) * common : more accurate sampling timing * eval-callback : minor fixes * cont : add time_meas impl * cont : fix log msg [no ci] * cont : fix multiple definitions of time_meas * llama-cli : exclude chat template init from time measurement * cont : print percentage of unaccounted time * cont : do not reset timings		2025-11-20 13:40:10 +02:00
..
batched-bench	batched-bench : add "separate text gen" mode (#17103 )	2025-11-10 12:59:29 +02:00
cvector-generator
export-lora
gguf-split
imatrix
llama-bench	bench : cache the llama_context state at computed depth (#16944 )	2025-11-07 21:23:11 +02:00
main	common : more accurate sampling timing (#17382 )	2025-11-20 13:40:10 +02:00
mtmd	mtmd-cli: Avoid logging to stdout for model loading messages in mtmd-cli (#17277 )	2025-11-15 12:41:16 +01:00
perplexity
quantize
rpc	Install rpc-server when GGML_RPC is ON. (#17149 )	2025-11-11 10:53:59 +00:00
run
server	webui: Add a "Continue" Action for Assistant Message (#16971 )	2025-11-19 14:39:50 +01:00
tokenize
tts
CMakeLists.txt