Hugging Face Daily Papers · · 5 min read

Ventor-QTest: Threat-Model-Driven Verification of Vendor-Hosted LLM APIs

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

As large language models become increasingly widespread, third-party providers that deploy open-weight models have become an important part of the ecosystem. Auditing the quality of their inference APIs is therefore an open problem. We formalize hosted model routing as a stochastic process and propose Ventor-QTest, a composite black-box audit that requires no probability information from the target API. Its repeated-request component sends each frozen constrained context to the target multiple times, reconstructs a categorical output distribution from the returned text counts, and reports average fidelity loss(AFL) as a null-bias-corrected, within-window mean coarsened-KL statistic. Its long-sequence component uses independent runs to report extreme fidelity loss (EFL) through the empirical upper tail of a run-level reference-centered-surprisal statistic. Across three logprob-capable route conditions, AFL shows strong linear descriptive agreement with a logprob-derived coarsened-KL comparator. Across seven route snapshots, 20-run sequence probes reveal route-specific EFL variation. AFL and EFL have little detectable route-level association with GPQA-Diamond accuracy. In contrast, pronounced EFL coincides with a decline in Terminal-Bench pass rate as task exposure increases. This pattern may arise because correctness in long-horizon tasks is more sensitive to extreme fidelity loss. These results motivate reporting AFL and EFL jointly, particularly when auditing long-horizon agentic tasks. The open-source implementation is available at \\url{<a href=\"https://github.com/Tencent/AI-Infra-Guard/tree/main/services/api_checker/ventor_qtest%7D\" rel=\"nofollow\">https://github.com/Tencent/AI-Infra-Guard/tree/main/services/api_checker/ventor_qtest}</a>.</p>\n","updatedAt":"2026-08-18T03:03:00.960Z","author":{"_id":"6a61f27af6b5a8f16f0277af","avatarUrl":"/avatars/197abee0edb281154f32fa8c4bee1029.svg","fullname":"wu","name":"xiangfanwu","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8998193144798279},"editors":["xiangfanwu"],"editorAvatarUrls":["/avatars/197abee0edb281154f32fa8c4bee1029.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.16391","authors":[{"_id":"6a83c932675db694db8cd4c4","name":"Xiangfan Wu","hidden":false},{"_id":"6a83c932675db694db8cd4c5","name":"Zonghao Ying","hidden":false},{"_id":"6a83c932675db694db8cd4c6","name":"Huiyu Wu","hidden":false},{"_id":"6a83c932675db694db8cd4c7","name":"Xing Zheng","hidden":false},{"_id":"6a83c932675db694db8cd4c8","name":"Huangsheng Cheng","hidden":false},{"_id":"6a83c932675db694db8cd4c9","name":"Xiaorong Shi","hidden":false},{"_id":"6a83c932675db694db8cd4ca","name":"Jing Guo","hidden":false}],"publishedAt":"2026-08-17T00:00:00.000Z","submittedOnDailyAt":"2026-08-18T00:00:00.000Z","title":"Ventor-QTest: Threat-Model-Driven Verification of Vendor-Hosted LLM APIs","submittedOnDailyBy":{"_id":"6a61f27af6b5a8f16f0277af","avatarUrl":"/avatars/197abee0edb281154f32fa8c4bee1029.svg","isPro":false,"fullname":"wu","user":"xiangfanwu","type":"user","name":"xiangfanwu"},"summary":"As large language models become increasingly widespread, third-party providers that deploy open-weight models have become an important part of the ecosystem. Auditing the quality of their inference APIs is therefore an open problem. We formalize hosted model routing as a stochastic process and propose \\textbf{Ventor-QTest}, a composite black-box audit that requires no probability information from the target API. Its repeated-request component sends each frozen constrained context to the target multiple times, reconstructs a categorical output distribution from the returned text counts, and reports average fidelity loss (AFL) as a null-bias-corrected, within-window mean coarsened-KL statistic. Its long-sequence component uses independent runs to report extreme fidelity loss (EFL) through the empirical upper tail of a run-level reference-centered-surprisal statistic. Across three logprob-capable route conditions, AFL shows strong linear descriptive agreement with a logprob-derived coarsened-KL comparator. Across seven route snapshots, 20-run sequence probes reveal route-specific EFL variation. AFL and EFL have little detectable route-level association with GPQA-Diamond accuracy. In contrast, pronounced EFL coincides with a decline in Terminal-Bench pass rate as task exposure increases. This pattern may arise because correctness in long-horizon tasks is more sensitive to extreme fidelity loss. These results motivate reporting AFL and EFL jointly, particularly when auditing long-horizon agentic tasks. The open-source implementation is available at https://github.com/Tencent/AI-Infra-Guard/tree/main/services/api_checker/ventor_qtest.","upvotes":12,"discussionId":"6a83c932675db694db8cd4cb","ai_summary":"Ventor-QTest audits hosted open-weight model APIs via repeated and long-sequence black-box probes, measuring average and extreme fidelity loss to detect degradation in long-horizon agentic performance.","ai_keywords":["Ventor-QTest","black-box audit","categorical output distribution","average fidelity loss","coarsened-KL","extreme fidelity loss","reference-centered-surprisal","GPQA-Diamond","Terminal-Bench","long-horizon agentic tasks"],"ai_summary_model":"thinkingmachines/Inkling-Small","organization":{"_id":"66543b6e420092799d2f625c","name":"tencent","fullname":"Tencent","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/5dd96eb166059660ed1ee413/Lp3m-XLpjQGwBItlvn69q.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"},{"_id":"66935bdc5489e4f73c76bc7b","avatarUrl":"/avatars/129d1e86bbaf764b507501f4feb177db.svg","isPro":false,"fullname":"Abidoye Aanuoluwapo","user":"Aanuoluwapo65","type":"user"},{"_id":"6a61f27af6b5a8f16f0277af","avatarUrl":"/avatars/197abee0edb281154f32fa8c4bee1029.svg","isPro":false,"fullname":"wu","user":"xiangfanwu","type":"user"},{"_id":"65372bae0d973d3fee4131c7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/OeBjr1hlUINmjUJ8iu42I.png","isPro":false,"fullname":"BadCat","user":"Foresta","type":"user"},{"_id":"6a6443f77fd42fcb33c15260","avatarUrl":"/avatars/79a72a7183cacbc3d9fe2ba1c297dab0.svg","isPro":false,"fullname":"Jiawei Hu","user":"jiawei2ch","type":"user"},{"_id":"67564ac4298969739a277049","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/HRIoTPmZjHUMCf87r2BXl.png","isPro":false,"fullname":"June","user":"June-Snow","type":"user"},{"_id":"62493535f1a5ac6c5d68a6ad","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1648964916160-noauth.png","isPro":false,"fullname":"Ch3nYe","user":"Ch3nYe","type":"user"},{"_id":"68e71969fa2b7fd74a3db481","avatarUrl":"/avatars/3c28937bbbfc6dd949fec8feb960d72d.svg","isPro":false,"fullname":"yingyuan pu","user":"sengle","type":"user"},{"_id":"6894eaf28ced68b510a4f00e","avatarUrl":"/avatars/a3f0b4f66f8cee0ba2043d1900144a03.svg","isPro":false,"fullname":"T4kiiiiiiiiiiiiii","user":"T4kiiiiiiii","type":"user"},{"_id":"6a840798b32fdc1131686a30","avatarUrl":"/avatars/d6e12246412f1de126e6c886e1ce29aa.svg","isPro":false,"fullname":"kellan","user":"kellanzz","type":"user"},{"_id":"6a8407bad06e65daf127249a","avatarUrl":"/avatars/91b4aa4660154d939dcead3d67ac3709.svg","isPro":false,"fullname":"DrKing","user":"DrKing0920","type":"user"},{"_id":"6750132f37c700d35223050a","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/RiuquTg7irzog4LFI24yA.png","isPro":false,"fullname":"xiangfanwu","user":"kexohs","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"66543b6e420092799d2f625c","name":"tencent","fullname":"Tencent","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/5dd96eb166059660ed1ee413/Lp3m-XLpjQGwBItlvn69q.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.16391.md","query":{}}">
Papers
arxiv:2608.16391

Ventor-QTest: Threat-Model-Driven Verification of Vendor-Hosted LLM APIs

Published on Aug 17
· Submitted by
wu
on Aug 18
Authors:
,

Abstract

Ventor-QTest audits hosted open-weight model APIs via repeated and long-sequence black-box probes, measuring average and extreme fidelity loss to detect degradation in long-horizon agentic performance.

As large language models become increasingly widespread, third-party providers that deploy open-weight models have become an important part of the ecosystem. Auditing the quality of their inference APIs is therefore an open problem. We formalize hosted model routing as a stochastic process and propose \textbf{Ventor-QTest}, a composite black-box audit that requires no probability information from the target API. Its repeated-request component sends each frozen constrained context to the target multiple times, reconstructs a categorical output distribution from the returned text counts, and reports average fidelity loss (AFL) as a null-bias-corrected, within-window mean coarsened-KL statistic. Its long-sequence component uses independent runs to report extreme fidelity loss (EFL) through the empirical upper tail of a run-level reference-centered-surprisal statistic. Across three logprob-capable route conditions, AFL shows strong linear descriptive agreement with a logprob-derived coarsened-KL comparator. Across seven route snapshots, 20-run sequence probes reveal route-specific EFL variation. AFL and EFL have little detectable route-level association with GPQA-Diamond accuracy. In contrast, pronounced EFL coincides with a decline in Terminal-Bench pass rate as task exposure increases. This pattern may arise because correctness in long-horizon tasks is more sensitive to extreme fidelity loss. These results motivate reporting AFL and EFL jointly, particularly when auditing long-horizon agentic tasks. The open-source implementation is available at https://github.com/Tencent/AI-Infra-Guard/tree/main/services/api_checker/ventor_qtest.

Community

Paper submitter about 6 hours ago

As large language models become increasingly widespread, third-party providers that deploy open-weight models have become an important part of the ecosystem. Auditing the quality of their inference APIs is therefore an open problem. We formalize hosted model routing as a stochastic process and propose Ventor-QTest, a composite black-box audit that requires no probability information from the target API. Its repeated-request component sends each frozen constrained context to the target multiple times, reconstructs a categorical output distribution from the returned text counts, and reports average fidelity loss(AFL) as a null-bias-corrected, within-window mean coarsened-KL statistic. Its long-sequence component uses independent runs to report extreme fidelity loss (EFL) through the empirical upper tail of a run-level reference-centered-surprisal statistic. Across three logprob-capable route conditions, AFL shows strong linear descriptive agreement with a logprob-derived coarsened-KL comparator. Across seven route snapshots, 20-run sequence probes reveal route-specific EFL variation. AFL and EFL have little detectable route-level association with GPQA-Diamond accuracy. In contrast, pronounced EFL coincides with a decline in Terminal-Bench pass rate as task exposure increases. This pattern may arise because correctness in long-horizon tasks is more sensitive to extreme fidelity loss. These results motivate reporting AFL and EFL jointly, particularly when auditing long-horizon agentic tasks. The open-source implementation is available at \url{https://github.com/Tencent/AI-Infra-Guard/tree/main/services/api_checker/ventor_qtest}.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.16391
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2608.16391 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2608.16391 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2608.16391 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers