News / #developer-tool Tag Developer Tool 500 articles archived under #developer-tool · RSS Sign in to follow Simon Willison community 25d ago PipeNetwork/minimax-h3-mlx PipeNetwork/minimax-h3-mlx MiniMax released MiniMax-H3 two days ago - they describe it as a "a general-purpose, omni-modal generative system", which in practice means it accepts text, images, audio and video and can use them to generate up to 15 second video clips with audio… 28 Simon Willison community 25d ago PipeNetwork/minimax-h3-mlx PipeNetwork/minimax-h3-mlx MiniMax released MiniMax-H3 two days ago - they describe it as a "a general-purpose, omni-modal generative system", which in practice means it accepts text, images, audio and video and can use them to generate up to 15 second video clips with audio… 26 GitHub Blog — AI & ML official-blog 25d ago How the GitHub legal team used Copilot CLI to streamline their workflows Learn how to build tools to simplify how you work—without writing a single line of code. The post How the GitHub legal team used Copilot CLI to streamline their workflows appeared first on The GitHub Blog . 19 Hugging Face Daily Papers research 25d ago Loud or Silent? A Reusable Framework for Per-Modality Failure Analysis in Multimodal Clinical AI Abstract Multimodal clinical models are usually judged on accuracy with every modality present, but deployment removes modalities; an echocardiogram is often unavailable where an ECG is routine. Two questions then matter beyond the size of the accuracy loss: which modality was… 14 MIT News — AI research 26d ago The benefits of medical AI assistance vary based on user expertise Study finds non-experts deferred to LLM-based diagnostic assistance, even when it was wrong, while clinicians caught AI errors. 22 arXiv — Machine Learning research 26d ago xMICD: Explainable Representation of Multiple ICD Codes arXiv:2608.00935v1 Announce Type: new Abstract: Electronic Health Records (EHRs) are widely used for clinical risk prediction using machine learning. International Classification of Diseases (ICD) codes provide structured information about patient diagnoses, but representing… 35 arXiv — Machine Learning research 26d ago Differentiable Lifting for Topological Neural Networks arXiv:2608.01160v1 Announce Type: new Abstract: Topological neural networks (TNNs) enable leveraging high-order structures on graphs (e.g., cycles and cliques) to boost the expressive power of message-passing neural networks. In turn, however, these structures are typically… 36 arXiv — Machine Learning research 26d ago Riemannian Attention Mechanisms for Transformers: A Theoretical Framework and Architecture Design arXiv:2608.01283v1 Announce Type: new Abstract: All Transformer-based large language models compute attention via the Euclidean inner product, an architectural choice that Dong et al. (2021) proved causes representational rank to decay doubly exponentially with depth in pure… 18 arXiv — Machine Learning research 26d ago Interpretable MEG Decoding of Perceived Speech: Cortical Sources and the Stimulus Features That Drive Retrieval arXiv:2608.01481v1 Announce Type: new Abstract: Short segments of perceived speech can be retrieved from non-invasive magnetoencephalographic (MEG) recordings by deep networks trained with a CLIP-style objective against wav2vec 2.0 audio embeddings. Yet their weights do not map… 38 arXiv — NLP / Computation & Language research 26d ago XL-DocBench: Benchmarking Evidence-Grounded Extra-Long Document Understanding arXiv:2608.00036v1 Announce Type: new Abstract: Real-world document tasks often ask professionals to answer questions from annual reports, regulations, clinical guidelines, and technical manuals that span hundreds or thousands of pages. Some questions also require comparing… 23 arXiv — NLP / Computation & Language research 26d ago MedUPS: Towards Diagnostic Assistance in Uncommon Medical Cases with Large Language Models arXiv:2608.01012v1 Announce Type: new Abstract: Uncommon and off-guideline cases are difficult for clinical decision support, because physicians must make a series of management decisions under diagnostic uncertainty and rarely see the full case at once. Most large language… 34 arXiv — NLP / Computation & Language research 26d ago Characterizing Treatment-Context Medication Evidence Across Clinic Notes and Structured EHR Medication History arXiv:2608.01570v1 Announce Type: new Abstract: Clinic notes and structured electronic health record (EHR) medication history often contain different medication information. Same-visit disagreement between these sources may result from note-side normalization errors, differences… 22 Vercel — AI dev-tools 26d ago Give your eve agent a browser Your eve agent can now navigate the web like a human with agent-browser . The @agent-browser/eve extension gives any eve agent a full set of browser tools: navigate pages, read content, click, fill forms, take screenshots, and inspect console and network activity. Everything… 33 Hacker News — AI on Front Page community 27d ago Andy Pavlo joins ClickHouse to establish ClickHouse Labs Article URL: https://clickhouse.com/blog/andy-pavlo-joins-clickhouse Comments URL: https://news.ycombinator.com/item?id=49156011 Points: 248 # Comments: 53 19 r/LocalLLaMA community 27d ago I compared MinerU, Granite-Docling, and PaddleOCR-VL on 12 PDF-parsing capabilities using 6 document types I tested them by sending the 6 documents, each meant to represent a different document type, through my own webapp and comparing every output against the source. All ran on the same L4 GPU. The documents: Financial statements with merged multi-level headers (A typical annual… 37 r/LocalLLaMA community 27d ago GLM 5.3 Spotted https://github.com/zai-org/z-ai-sdk-java/commits/glm-5.3   submitted by   /u/Few_Painter_5588 [link]   [comments] 10 arXiv — Machine Learning research 27d ago Technological Advances in Detecting and Managing Cognitive Impairment in Older Adults: Trends, Challenges, and Future Directions arXiv:2607.28687v1 Announce Type: new Abstract: As populations age, cognitive decline from mild cognitive impairment (MCI) to dementia is a defining health challenge of the coming decades, yet routine assessment often misses its earliest signs. This article critically… 30 arXiv — Machine Learning research 27d ago Mitigating Class-Tail Undercoverage in Medical Vision-Language Models under Clinical Shift arXiv:2607.28696v1 Announce Type: new Abstract: Medical vision-language models (VLMs) can retain high observed marginal coverage after clinical shift while substantially under-covering an individual disease class. The affected class varies with acquisition protocol and backbone… 9 arXiv — Machine Learning research 27d ago MMFGU: Multimodal Federated Graph Unlearning arXiv:2607.28708v1 Announce Type: new Abstract: Multimodal federated graph learning enables clients to collaboratively train graph models over structural, textual, and visual signals without sharing private local data. However, the presence of heterogeneous multimodal content… 32 arXiv — Machine Learning research 27d ago Federated Foundation Models Fine-Tuning with Heterogeneous Compressed Clients arXiv:2607.29071v1 Announce Type: new Abstract: Federated learning of foundation models faces a fundamental resource-asymmetry challenge: the institutions holding the most valuable domain-specific data cannot host billion-parameter models. Existing heterogeneous federated… 27 arXiv — Machine Learning research 27d ago PiDDM: Physics-Informed Differentiable Degradation Modeling for Lithium-Ion Battery State-of-Health Prediction arXiv:2607.29095v1 Announce Type: new Abstract: Accurate prediction of lithium-ion battery state of health (SOH) is essential for reliable energy storage operation. However, purely data-driven models may generalize poorly across cycling protocols and produce physically… 27 arXiv — Machine Learning research 27d ago StraightDP: Geometry-Aware Differential Privacy for Rectified-Flow Transformers arXiv:2607.29100v1 Announce Type: new Abstract: Differentially private (DP) training of text-conditioned generative models suffers a utility cliff at strong privacy. We revisit this problem through the geometry of rectified flows: along the straight interpolation between noise… 12 arXiv — Machine Learning research 27d ago Fracture Risk Prediction in Adults Over 50 Years Old Using DXA and EHR: Comparison of Traditional and Machine Learning Models in Two Large Cohorts arXiv:2607.28671v1 Announce Type: cross Abstract: Accurate fracture risk prediction is important for osteoporosis management, but commonly used clinical tools may not fully use information available in electronic health records (EHRs) and dual-energy X-ray absorptiometry (DXA)… 31 arXiv — NLP / Computation & Language research 27d ago Rolling With Resistance: Preference-Optimized LLM Counselors Can Trade Goal Persistence for Relational Attunement in Motivational Interviewing arXiv:2607.28814v1 Announce Type: new Abstract: In Motivational Interviewing (MI), a client's sustain talk (arguments for the status quo) calls for the counselor to roll with resistance, a move that can fail in two opposite ways: capitulation (abandoning the change agenda to… 38 arXiv — NLP / Computation & Language research 27d ago FriendBench: Benchmarking Dyadic Familiarity Inference in Humans and Multimodal Large Language Models arXiv:2607.29602v1 Announce Type: new Abstract: Reading a social situation often depends on behavior, not words alone. We introduce FriendBench, a benchmark for inferring whether two people are already familiar or are meeting as strangers, from a 20-second clip of a dyadic… 5 llama.cpp releases dev-tools 28d ago b10219 cli : persist reasoning_content in chat history ( #26362 ) cli : persist reasoning_content in chat history llama-cli collected reasoning from the stream for display but only stored assistant content in messages, so --reasoning-preserve could not re-inject prior thoughts on later… 22 r/MachineLearning community 29d ago VLMs can score well on benchmarks, while silently erasing meaningful terms and including hallucinate bias [P] While working with VLMs for report generation on chest x-rays (RRG), we noticed that evaluation metrics are flawed. Flawed in a sense where they rewarded repetitive templates, reports without clinical terms and reports which were "normal" with high scores on benchmark metrics.… 24 Simon Willison community 29d ago llm-mcp-client 0.1a0 Release: llm-mcp-client 0.1a0 See this blog entry . Tags: llm , model-context-protocol 5 llama.cpp releases dev-tools 29d ago b10211 vulkan: update vulkan sdk to 1.4.357.0 ( #26303 ) Website: https://llama.app macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64… 27 NVIDIA Developer Blog official-blog 29d ago NVIDIA Video Codec SDK 13.1: Zero-Copy Transcode, AV1 B-Frames, and Frame-Accurate Seek The demand for high-quality video continues to accelerate across industries, powering everything from immersive streaming experiences to remote collaboration,... 29 OpenAI Python SDK releases dev-tools 29d ago v2.52.0 2.52.0 (2026-07-31) Full Changelog: v2.51.0...v2.52.0 Features api: content provenance checks ( 1d6c118 ) Bug Fixes client: honor Retry-After delays up to two minutes ( #3555 ) ( 7fa7946 ) Documentation add API-key mTLS HTTP client recipes ( #3552 ) ( 7a3d5e4 ) 30 arXiv — Machine Learning research 1mo ago Position, Not Provenance: Separating Reasoning Mediation from Sycophancy in Medical Vision-Language Models arXiv:2607.27304v1 Announce Type: new Abstract: Medical vision-language models (VLMs) generate chain-of-thought (CoT) reasoning before answering clinical questions, but whether this reasoning causally influences predictions remains unclear. We present CoT-Mediate, a behavioral… 14 arXiv — Machine Learning research 1mo ago ECG-InterpBench: Benchmarking the Interpretability of ECG Foundation Models with Matched-Scale Sparse Autoencoders arXiv:2607.27404v1 Announce Type: new Abstract: Existing benchmarks for electrocardiogram foundation models primarily evaluate downstream predictive performance, providing limited insight into whether their internal representations can be faithfully decomposed, clinically… 8 arXiv — Machine Learning research 1mo ago FedOGL: Combating Catastrophic Forgetting in Federated Open-World Multimodal Graph Learning arXiv:2607.27665v1 Announce Type: new Abstract: Federated graph learning enables collaborative training over decentralized graph data without sharing raw graph information. As such risks evolve, clients must learn emerging classes from private multimodal graph streams, retain… 26 arXiv — Machine Learning research 1mo ago LoRA Scaffolded Policy Optimization (LSPO): A Sampling-Time Low-Rank Scaffold for Recovering Reinforcement-Learning Gradient on Zero-Reward Cliff Prompts arXiv:2607.27787v1 Announce Type: new Abstract: Reinforcement learning from verifiable rewards (RLVR) for mathematical reasoning suffers from a structural blind spot: on "cliff" prompts-those on which every sampled rollout in a group fails-the group-normalized advantage is… 4 arXiv — Machine Learning research 1mo ago Encryption-Compatible Clustered Federated Learning via Distributed Expectation-Maximization over Metadata arXiv:2607.28338v1 Announce Type: new Abstract: Clustered Federated Learning (CFL) addresses data heterogeneity in federated settings by grouping clients with similar data distributions to enable effective training. Existing methods face a trade-off between privacy preservation,… 15 arXiv — NLP / Computation & Language research 1mo ago Same Facts, Different Diagnosis: Measuring and Mitigating Narrative Anchoring in Clinical Language Models arXiv:2607.27384v1 Announce Type: new Abstract: Large language models used for clinical diagnostic reasoning are sensitive to sociolinguistic register, not just clinical content. We term this failure mode Narrative Anchoring: identical clinical facts expressed in different… 26 arXiv — NLP / Computation & Language research 1mo ago The MADRS Pipeline: Supporting Depression Assessment in Clinical Trials arXiv:2607.28190v1 Announce Type: new Abstract: Depression is a major mental disorder for which diagnosis relies primarily on clinical assessments. Automated methods to support its detection via the psychiatric MADRS scale are getting more and more attention. While existing… 10 arXiv — NLP / Computation & Language research 1mo ago Digital Harf: A Clinically Integrated Multimodal AI System for Pervasive Arabic Speech and Language Therapy arXiv:2607.27212v1 Announce Type: cross Abstract: Children with Autism Spectrum Disorder in Arabic-speaking countries face compounded barriers to effective speech and language therapy: a shortage of qualified specialists, limited service reach beyond urban centers, and a… 26 Hugging Face Daily Papers research 1mo ago Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents Abstract GUI agents have the potential to become a general purpose executor over existing digital devices. To advance them toward real-world use, we envision agents that operate reliably on real devices, execute workflows across platforms, combine GUI interaction with CLI… 31 Hugging Face Daily Papers research 1mo ago SpatialCLI: Learning to Reason With Spatial Tools, Then Without Them Abstract Vision-language models (VLMs) are increasingly used in embodied agents to interpret visual inputs, reason about spatial relationships, and make task-level decisions based on that reasoning. However, a fundamental capability mismatch remains: general VLMs can reason… 14 Vercel — AI dev-tools 1mo ago Vercel MCP now supports the 2026-07-28 MCP specification Vercel MCP now supports the 2026-07-28 MCP specification, giving newer clients a stateless request model and updated authorization behavior without any change on the client side. Clients built for the 2025 protocol keep working exactly as before, and clients that understand the… 36 Vercel — AI dev-tools 1mo ago Chat SDK now supports reactions and ephemeral messages on Teams Chat SDK's Microsoft Teams adapter now supports reactions and ephemeral messages through the same API as other adapters. Bots can add and remove reactions on Teams messages, and call thread.postEphemeral() or channel.postEphemeral() to send native, targeted messages that only… 5 r/LocalLLaMA community 1mo ago Smallest model (& tips) for intelligent computer use via Hermes? Hello, I have a friend who's using various local LLM's like qwen3.6 27B, 35b-a3b, North Mini Code, and qwen2.5-vl-7b (just for vision). They have a use case where they're trying to have an LLM drive an actual machine via hermes' computer_use tool and cua_driver to click through… 7 Vercel — AI dev-tools 1mo ago Server-Timing response headers will pass through to the client Vercel's CDN will begin passing through the Server-Timing response header to the client on August 10, 2026. Use Server-Timing to report backend metrics like database query time and cache hits. These values appear in the browser's network panel and as PerformanceServerTiming… 11 Hugging Face Daily Papers research 1mo ago SecRespond: Benchmarking AI Agents for Real-World Post-Compromise Incident Response Abstract Large Language Model (LLM) agents are increasingly adopted in real-world security operations with access to host artifacts and command-line interfaces (CLIs), making it critical to thoroughly assess their security capabilities. However, existing cybersecurity benchmarks… 36 arXiv — NLP / Computation & Language research 1mo ago FedWeave: Rethinking the Unit of Specialization in Heterogeneous Federated MoE-LoRA arXiv:2607.26618v1 Announce Type: cross Abstract: Federated PEFT enables LLMs to collaboratively adapt to decentralized private data without sharing raw examples. However, task heterogeneity across clients can cause cross-task interference and gradient conflicts during… 7 arXiv — NLP / Computation & Language research 1mo ago SecRespond: Benchmarking AI Agents for Real-World Post-Compromise Incident Response arXiv:2607.26791v1 Announce Type: cross Abstract: Large Language Model (LLM) agents are increasingly adopted in real-world security operations with access to host artifacts and command-line interfaces (CLIs), making it critical to thoroughly assess their security capabilities.… 32 arXiv — NLP / Computation & Language research 1mo ago $\texttt{AMEND++}$: Benchmarking Eligibility Criteria Amendments in Clinical Trials arXiv:2601.06300v2 Announce Type: replace Abstract: Clinical trial amendments frequently introduce delays, increased costs, and administrative burden, with eligibility criteria being the most commonly amended component. We introduce \textit{eligibility criteria amendment… 15 Vercel — AI dev-tools 1mo ago Run multiple isolated agents in a single Sandbox The @vercel/sandbox SDK now supports multiple Linux users and groups, so you can run agents side by side in a single Sandbox. Each agent runs as its own user with a private home directory. A group opens a shared workspace when they need to collaborate. This makes multi-agent… 14 Page 6 of 10 · 500 articles ← Newer Older →