llama_cpp_for_radxa_dragon_wing_q6a

pingu_98/llama_cpp_for_radxa_dragon_wing_q6a

History

Max Krasnyansky 053b1539c0 threading: support for GGML_SCHED_PRIO_LOW, update thread info on Windows to avoid throttling (#12995 ) * threading: support for GGML_SCHED_PRIO_LOW, update thread info on Windows to avoid throttling We talked about adding LOW priority for GGML threads in the original threadpool PR. It might be useful for some cases to avoid contention. Latest Windows ARM64 releases started parking (offlining) the CPU cores more aggresively which results in suboptimal performance with n_threads > 4. To deal with that we now disable Power Throttling for our threads for the NORMAL and higher priorities. Co-authored-by: Diego Devesa <slarengh@gmail.com> * threading: disable SetThreadInfo() calls for older Windows versions * Update tools/llama-bench/llama-bench.cpp Co-authored-by: Diego Devesa <slarengh@gmail.com> --------- Co-authored-by: Diego Devesa <slarengh@gmail.com>		2025-05-31 15:39:19 -07:00
..
cmake
arg.cpp	threading: support for GGML_SCHED_PRIO_LOW, update thread info on Windows to avoid throttling (#12995 )	2025-05-31 15:39:19 -07:00
arg.h
base64.hpp
build-info.cpp.in
chat-parser.cpp	server: allow unclosed thinking tags (#13931 )	2025-05-31 08:26:10 -07:00
chat-parser.h	sync : vendor (#13901 )	2025-05-30 16:25:45 +03:00
chat.cpp	sync : vendor (#13901 )	2025-05-30 16:25:45 +03:00
chat.h	server: fix streaming crashes (#13786 )	2025-05-26 16:03:57 +01:00
CMakeLists.txt	sync : vendor (#13901 )	2025-05-30 16:25:45 +03:00
common.cpp	threading: support for GGML_SCHED_PRIO_LOW, update thread info on Windows to avoid throttling (#12995 )	2025-05-31 15:39:19 -07:00
common.h	server: --offline mode (#13804 )	2025-05-26 22:34:27 +01:00
console.cpp
console.h
json-partial.cpp	sync : vendor (#13901 )	2025-05-30 16:25:45 +03:00
json-partial.h	sync : vendor (#13901 )	2025-05-30 16:25:45 +03:00
json-schema-to-grammar.cpp	sync : vendor (#13901 )	2025-05-30 16:25:45 +03:00
json-schema-to-grammar.h	sync : vendor (#13901 )	2025-05-30 16:25:45 +03:00
llguidance.cpp
log.cpp
log.h
ngram-cache.cpp
ngram-cache.h
regex-partial.cpp
regex-partial.h
sampling.cpp	`server`: streaming of tool calls and thoughts when `--jinja` is on (#12379 )	2025-05-25 01:48:08 +01:00
sampling.h
speculative.cpp
speculative.h