Skip to content
View dolev31's full-sized avatar

Highlights

  • Pro

Block or report dolev31

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
dolev31/README.md

Ido Levy

Research on LLM agents at IBM and the Weizmann Institute of Science:
what agents pursue without being asked, and how safely they act.

Google Scholar LinkedIn Hugging Face New paper


🔎 Asking for What Was Never Requested: Horizontal and Vertical Proactivity in Agents · 2026

Ido Levy, Asaf Yehudai, Segev Shlomov, Asaf Adi, Leshem Choshen

Asking for What Was Never Requested: one customer request handled by the same model, prompted and trained with Q&D

What should an LLM agent pursue that the user never asked for? Need graphs measure it without a model judge, and Q&D trains it from the consequences of its own questions. The trained 8B questioner recovers 90% of the required evidence where the same model, prompted, recovers 78%, and it more than doubles retail task success in a τ²-bench customer-service agent it was never trained on.

Stars

Paper · Project page · Code · Model 🤗

🛡️ ST-WebAgentBench: A Benchmark for Evaluating Safety and Trustworthiness in Web Agents · 2024

Ido Levy, Ben Wiesel, Sami Marreed, Alon Oved, Avi Yaeli, Segev Shlomov

A benchmark for evaluating the safety and trustworthiness of web agents in enterprise scenarios.

Stars

Paper · Website · Code · Leaderboard · Dataset

🧩 CUGA: an open-source generalist agent for the enterprise · contributor

Stars

An agent harness for complex tasks on the web and APIs, with OpenAPI and MCP integrations, a composable architecture, reasoning modes and policy-aware features.

Website · Code

Pinned Loading

  1. ProactiveInquirer ProactiveInquirer Public

    Asking for What Was Never Requested: horizontal and vertical proactivity in LLM agents. Need-graph metrics (no LLM judge) and Q&D, which trains a questioner from the consequences of its questions.

    Python 9

  2. cuga-project/cuga-agent cuga-project/cuga-agent Public

    CUGA is an open-source generalist agent harness for the enterprise, supporting complex task execution on web and APIs, OpenAPI/MCP integrations, composable architecture, reasoning modes, and policy…

    Python 882 162

  3. segev-shlomov/ST-WebAgentBench segev-shlomov/ST-WebAgentBench Public

    A Benchmark for Evaluating Safety and Trustworthiness in Web Agents for Enterprise Scenarios

    Python 29 8