llama_cpp_for_radxa_dragon_.../tools
Georgi Gerganov 196f5083ef
common : more accurate sampling timing (#17382)
* common : more accurate sampling timing

* eval-callback : minor fixes

* cont : add time_meas impl

* cont : fix log msg [no ci]

* cont : fix multiple definitions of time_meas

* llama-cli : exclude chat template init from time measurement

* cont : print percentage of unaccounted time

* cont : do not reset timings
2025-11-20 13:40:10 +02:00
..
batched-bench
cvector-generator
export-lora
gguf-split
imatrix
llama-bench
main common : more accurate sampling timing (#17382) 2025-11-20 13:40:10 +02:00
mtmd mtmd-cli: Avoid logging to stdout for model loading messages in mtmd-cli (#17277) 2025-11-15 12:41:16 +01:00
perplexity
quantize
rpc
run
server webui: Add a "Continue" Action for Assistant Message (#16971) 2025-11-19 14:39:50 +01:00
tokenize
tts
CMakeLists.txt