Tracked Repositories
1310
Cognaptus DataHub Monitor
A monitored reference page for GitHub repositories surfaced from arXiv-paper digests, rendered from a machine-generated local data file.
Tracked Repositories
1310
Unique Papers
935
Core Fields
paper title, ref_id, GitHub URL
Refresh Mode
local data file written by backend tasks
Use it as a lightweight index of implementation assets surfaced from research digests. The page stays defensive: required fields remain visible even when optional metadata is missing.
This page is designed as a refreshable reference surface rather than a hand-maintained article.
The goal is simple: when paper-digest workflows identify linked GitHub repositories, keep them visible in one place with enough context to scan quickly and revisit later.
Search by paper title, arXiv reference, GitHub repository, author, or tags when those fields are available.
Repository containing code, configurations, prompts, and data-processing utilities for according-to prompting and QUIP-Score experiments.
Repository containing code and results for the LLM-based zero-shot tree induction and embedding experiments, with feature-description files referenced for private-dataset feature names.
Repository containing code and data for replicating the paper's LIME and SP-LIME experiments.
Repository linked by the paper as the code for PriorDynaFlow, the proposed a priori dynamic multi-agent workflow construction framework.
Repository for the KindsOfReasoning collection and raw outputs/evaluation results for OpenAI instruction-tuned models.
Repository identified by the paper as containing code for the empirical experiments.
Repository containing the experimental software associated with the proposed causal evaluation framework for deferring systems.
Python, notebook, and spreadsheet materials for driver ranking, forced ARMA changes, and related paper experiments.
EvolveGCN repository data folder referenced as the source for the SBM dynamic graph benchmark.
Repository containing datasets and code files supporting the paper's DeepSeek-versus-other-LLMs benchmark.
Repository containing accompanying code and resources for the book, including chapter folders and Python-oriented examples for XAI techniques.
An implementation referenced for reducing GPU all-to-all communication bottlenecks in large-scale MoE execution.
A GitHub repository associated with the survey that organises resources and representative work on self-evolving AI agents.
Repository containing the EIIE portfolio-management code, configurations, figures, tables, and data used for the cryptocurrency replication and stock-market extension.
AIDE-ML: AI-Driven Exploration in the Space of Code, used as the agentic AutoML baseline/interface for observing decisions and artifacts.
Git repository providing open access to the research datasets supporting the study's findings and replication.
Source code for a minimal MedLog prototype using an OpenAPI-described HTTP REST interface.
Repository for the financial news dataset used with permission as part of the paper's Bloomberg dataset experiments.
Open implementation code for the TGN-SEAL framework introduced and evaluated in the paper.
Public codebase containing the Agent-Driver pipeline, tool library, cognitive memory, reasoning engine, scripts for fine-tuning and inference, prepared data layout, and nuScenes evaluation workflow.
The public Lean 4/Mathlib library containing the formalized QAOA and Ising-ring components, the lower-bound and attainability theorems, and the machine-checked proof of residualEnergy_isLeast.
Repository released by the authors for the minimal agentic theorem prover and experiment reproduction.
Repository for the CLINC150 intent-classification and out-of-scope evaluation data used in the experiments.
Repository for the StackOverflow short-text intent dataset evaluated by the paper.
Repository containing the Banking77 fine-grained banking-intent dataset evaluated by the paper.
Repository for a hybrid LSTM, multi-head attention, Gaussian fuzzy-rule, and ARIX local-model forecasting system, including folders for LSTM, Transformer, Neuro-fuzzy, data utilities, model code, and an experiment notebook.
R-based analysis and reporting pipeline orchestrated with DVC, with scripts, parameters, environment configuration, intermediate outputs, and instructions for reproducing the reported results.
Repository containing code used for the paper, including notebooks and utilities for the daily trading strategy, the CNN model with the new loss, data downloading/processing, baselines, and analysis.
Repository containing Python, R, and notebook implementations, data files, portfolio outputs, and figures associated with the dynamic stock-recommendation study.
A repository created to keep pace with the fast-moving literature on LLMs and LLM-based agents in science.
Repository released by the author for resources associated with the LLM-agent survey.
TensorTrade is cited as an open-source package/platform relevant to implementing and adapting RL methods to finance.
Official code repository for the paper's Agora implementation and heterogeneous 100-agent demonstration.
The GitHub repository contains the self-improving coding-agent implementation introduced by the paper.
FakeNewsNet repository containing fake-news research data resources and crawler tooling for collecting news and related social-media data.
Repository containing the resulting mind map, collected notes for each reviewed grey resource, and an online repository of references for further exploration.
The Google Agent2Agent protocol repository, reviewed as a general-purpose inter-agent protocol and used in the paper's comparative use-case analysis.
Repository for the Web-Agent Protocol, reviewed as a domain-specific inter-agent or human-computer interaction protocol.
Repository for the agents.json specification, reviewed as a domain-specific context-oriented protocol for exposing website capabilities to agents.
Companion repository maintained by the authors to track ongoing developments in AI agent protocols.
An Awesome Data Agents repository linked directly in the paper header, likely used to collect or organize data-agent resources associated with the survey.
JoyAgent is discussed as a proto-L3 system that begins to address predefined-toolset limitations through tool evolution and multi-level thinking.
GitHub repository established by the authors as the project page associated with the survey on embodied learning for object-centric robotic manipulation.
Repository linked by the paper as the full list of surveyed papers and summary slides.
Repository for AlphaFin, a benchmark/resource associated with financial question answering and stock prediction.
Repository for R-Judge, a benchmark for safety judgment and risk identification.
Repository for FinanceBench, a financial question-answering benchmark.
Repository for a Japanese financial language-model benchmark.
Repository for BBT-Fin/CUGE-related Chinese financial language benchmark resources.
Repository for FinEval, a Chinese benchmark for financial domain knowledge.
Repository for PIXIU/FinMA-related financial LLM resources, instruction data, and evaluation benchmarks.
Repository for CFBenchmark, a Chinese financial benchmark covering multiple financial NLP tasks.
Repository for DocMath-Eval, used for evaluating numerical reasoning over text and tables.
A curated repository of papers and benchmarks associated with the survey on reasoning with foundation models.
agentUniverse, a multi-agent ecosystem for autonomous agents.
Agno, described as an agentic workflow framework for LLM applications.
Phidata, described as a framework for building multi-modal agents with memory, knowledge, tools, and reasoning.
Coze, described as an open framework for building agentic applications.
Flowise, described as a drag-and-drop UI to build LLM apps with LangChain.
LangGraph, described as a stateful multi-actor workflow library for LLM applications.
Dify, described as an open-source LLM application development platform.
Microsoft Semantic Kernel, compared as an agent workflow system.
n8n, described as a fair-code workflow automation platform with UI and integrations.
OpenAI Swarm, described as a multi-agent framework by OpenAI.
Qwen-Agent, a QwenLM agent framework/repository compared in the survey.
An author-linked repository organizing papers on retrieval-augmented generation.
GPT Engineer, a software-development agent implementation cited in the engineering application survey and open-source project discussion.
GPT Researcher, an experimental application that uses LLMs for research-question development, web crawling, source summarization, and aggregation.
AI Legion, an LLM-agent implementation cited in the survey's open-source library and reference set.
LoopGPT, an LLM-agent implementation cited in the survey's open-source library and reference set.
AGiXT, an agent framework implementation cited in the survey as a dynamic AI automation platform.
DemoGPT, a software-development agent repository cited in the engineering application survey and open-source project discussion.
MiniAGI, an LLM-agent implementation cited in the survey's open-source library and reference set.
AgentVerse, a multi-agent collaboration framework referenced among surveyed agent systems and open-source libraries.
AgentGPT, an LLM-based autonomous-agent system cited in the survey's open-source library and reference set.
Auto-GPT, an autonomous LLM-agent implementation included in the construction taxonomy and open-source library discussion.
SmolModels/developer-style agent repository cited as a software engineering application artifact.
WorkGPT, a workflow-oriented LLM-agent framework cited as similar to AutoGPT and LangChain.
SuperAGI, an autonomous-agent framework cited in the survey's open-source library and reference set.
XLang, an LLM-agent/tool-use framework cited as supporting executable language grounding and interaction with databases, web applications, and physical robots.
Repository for the survey on memory mechanisms of LLM-based agents, including the paper link and visual summaries of the survey sections.
A curated reading list accompanying the survey, organized around LLM-agent optimization methods, datasets, benchmarks, and applications.
Repository for the paper's code and experimental workflow comparing LLM self-explanations, human rationales, and post-hoc attribution explanations.
Repository path listed as the source for the gas station revenue dataset.
Repository listed as the source for daily COVID-19 confirmed and recovered case data.
Repository listed as the source for SPMD and VED driving and vehicle energy datasets.
Repository listed as the source for the exchange-rate dataset used in LTSF studies.
Repository listed as the source for daily stock opening price data.
Repository listed as the source for the ETT transformer temperature/load dataset.
Repository containing the medical RAG application, evaluation framework, text rechunking pipeline, configurations, and generated evaluation outputs.
OpenAI HumanEval benchmark for evaluating code generation with pass@k metrics.
Self-rewarding reasoning LLM implementation referenced as an example of using model-generated judgments to reduce annotation cost.
AlpacaEval benchmark for automatic evaluation of instruction-following model outputs.
Repository containing the Phase 1 SSL and contrastive encoder code, Phase 2 MoE PPO curriculum, inference and backtest scripts, routing diagnostics, and deployment-related components; the repository states that Phase 3 personalization is proprietary and not released.
Repository containing code for training and evaluating the Transformer electricity price forecasting model and comparing it against EPF toolbox benchmarks.
Repository linked from the paper that mirrors the paper title, abstract, framework figure, table of contents, and compiled relevant works for agent categories and attack types.
Author-provided production-oriented implementation of the A-Mem agentic memory system.
Author-provided repository for evaluating the A-Mem method and reproducing benchmark experiments.
Repository for ABC-Bench, a benchmark for evaluating whether coding agents can explore repositories, edit code, configure environments, deploy containerized backend services, and pass external HTTP/API integration tests.
Optimized Faiss exact-search implementation for Intel CPUs.
Cycle-approximate IKS simulator parameterized with RTL and memory/interconnect timing.
Public repository containing the access-controlled website implementation, agent-side experiment scripts, modified SST/Auth component as a submodule, and experiment logs for delegated access workflows.
Code and dataset repository for ActivationReasoning, including implementation and materials for replication or extension.
Official repository containing the FLARE code, task configurations, experimental data, retrieval setup, and run instructions.
Repository containing data loaders, preprocessing, regime models, custom environments, PPO training pipelines, ablations, evaluation outputs, figures, and notebooks associated with the regime-aware portfolio framework.
SentiFin is cited as a benchmark dataset for sentiment analysis of Indian financial news headlines and is used to fine-tune the LLaMA 3.2 model.
An installable Python repository implementing AGoT, AIoT, and GIoT and providing experimental setups, datasets, and result files used to reproduce or inspect the paper's evaluations.
Repository containing the code developed and used for the study, with a README to support replication and methodology exploration.
GitHub repository identified by the paper as containing implementation details for the Dueling DDQN liquidity-provision method and baseline methods.
Repository stated to include scripts for data preprocessing, model training, performance evaluation, and ablation studies for AMDTL.
Repository containing AMDM implementation code, simulation scripts, evaluation scripts, example data, results files, plots, and the paper materials.
Longformer repository for the long-document transformer model evaluated as an additional architecture in the paper.
A continuously updated collection of FFM-related publications, tools, datasets, and resources associated with the survey.
The paper identifies this repository as containing representative Python code generated under the tested prompt levels.
Claude-Agent-SDK framework used as one of the scaffold alternatives in the agentic scaffold impact study.
Public repository for the AgencyBench benchmark and evaluation toolkit released by the authors.
OpenAI-Agents-SDK framework used as one of the scaffold alternatives in the agentic scaffold impact study.
Repository for the AGENT KB cross-framework agent memory system introduced and evaluated in the paper.
Repository for the Agent Mentor / Agent Analytics open-source observability and analytics platform for agentic AI applications.
GitHub path identified by the paper as the code corresponding to the analytics pipeline used for semantic feature analysis.
Repository containing experiment code, response matrices, IRT models, task embeddings, LLM-as-a-judge features, adaptive testing code, and scripts for the new task, new response, new agent, and new benchmark experiments.
Repository for the paper's virtual trading arena, ArenaTrader implementation, prompts, code, and data.
Official code repository implementing AWM pipelines for WebArena and Mind2Web in offline and online settings.
Repository for the Agent-as-a-Judge project and DevAI-related evaluation artifacts.
GPT-Pilot is one of the three open-source code-generation agentic systems benchmarked in the paper.
A GitHub repository collecting papers and resources for the survey on Agent-as-a-Judge.
Repository associated with Agent-R, the paper's iterative self-training framework for training language-model agents to reflect and recover from errors.
Repository for AgentBench datasets, environments, and integrated evaluation package.
Indeed Hiring Lab repository tracking the share of job postings mentioning artificial intelligence.
Google's Agent-to-Agent Protocol repository, referenced as the source for A2A, one of the modern agent communication protocols compared in the paper.
BlenderMCP repository cited as a GitHub integration for Blender Model Context Protocol.
PowerAgent PowerMCP repository cited as a GitHub implementation for power-system simulation software MCPs.
Repository hosting the draft proteomics_GROUNDING.md epistemic grounding specification and Appendix A preliminary testing materials.
LangGraph is used as an example of developer-defined graph/state-machine execution with persistence, checkpoints, controlled cycles, guard nodes, and approvals.
Swarm is used as an example of star-with-handoffs orchestration using lightweight specialists, routines, and controller selection.
Agent-zero repository evaluated as a selected agentic framework in the paper's 22-framework benchmark study.
ANUS repository evaluated as a selected agentic framework in the paper's 22-framework benchmark study.
Camel repository evaluated as a selected agentic framework in the paper's 22-framework benchmark study.
CrewAI repository evaluated as a selected agentic framework in the paper's 22-framework benchmark study.
MetaGPT repository evaluated as a selected agentic framework in the paper's 22-framework benchmark study.
Google ADK repository evaluated as a selected agentic framework in the paper's 22-framework benchmark study.
Source code repository released by the authors for reproducing the benchmark comparison.
LangChain repository evaluated as a selected agentic framework in the paper's 22-framework benchmark study.
LangGraph repository evaluated as a selected agentic framework in the paper's 22-framework benchmark study.
Mastra repository evaluated as a selected agentic framework in the paper's 22-framework benchmark study.
PraisonAI repository evaluated as a selected agentic framework in the paper's 22-framework benchmark study.
Autogen repository evaluated as a selected agentic framework in the paper's 22-framework benchmark study.
Semantic-kernel repository evaluated as a selected agentic framework in the paper's 22-framework benchmark study.
TaskWeaver repository evaluated as a selected agentic framework in the paper's 22-framework benchmark study.
OpenAI-Agents-Python repository evaluated as a selected agentic framework in the paper's 22-framework benchmark study.
Swarm repository evaluated as a selected agentic framework in the paper's 22-framework benchmark study.
Pydantic-AI repository evaluated as a selected agentic framework in the paper's 22-framework benchmark study.
Qwen-Agent repository evaluated as a selected agentic framework in the paper's 22-framework benchmark study.
AutoGPT repository evaluated as a selected agentic framework in the paper's 22-framework benchmark study.
SuperAGI repository evaluated as a selected agentic framework in the paper's 22-framework benchmark study.
Upsonic repository evaluated as a selected agentic framework in the paper's 22-framework benchmark study.
Agency-swarm repository evaluated as a selected agentic framework in the paper's 22-framework benchmark study.
BabyAGI repository evaluated as a selected agentic framework in the paper's 22-framework benchmark study.
A curated and updateable repository of papers and resources organized according to the survey's agentic reasoning taxonomy.
Code repository for the Agentic Reasoning framework introduced and evaluated in the paper.
Public repository for Agentic Reward Modeling, including the RewardAgent implementation and materials intended to facilitate further research.
A continuously updated collection of relevant studies for the Agentic Web.
Repository for the AgenticPay benchmark and framework, including buyer and seller agents, negotiation environments, examples, metrics, and model backends for LLM-based commerce negotiation.
Repository containing the AgentLAB benchmark code, attack scripts, environments, prompts, data, and usage instructions for evaluating LLM agents against long-horizon attacks.
Repository for the AgentRewardBench library, including tools for running agents, running judges, scoring judgments, loading trajectories, and submitting leaderboard results.
Repository associated with AgentRx, the paper's diagnostic framework and benchmark artifact for AI-agent failure attribution.
Official repository for PASB, including the benchmark data, baseline runners, judge code, audit scripts, and documentation.
Hermes-Agent is one of the two stateful personal-agent stacks evaluated by PASB.
OpenClaw is one of the two stateful personal-agent stacks evaluated by PASB.
A curated GitHub repository associated with the paper that lists and classifies related papers on LLM-based agents in software engineering.
The GitHub repository for AgentScope, the multi-agent platform described in the paper.
Public implementation of AgentStepper and relevant data for the paper's evaluation.
GitHub Gist containing the Claude Code implementation prompt for the case summarization by file name microservice.
GitHub Gist containing the Case Summarization by Given Case Name Workflow pitch generated by the Planning Agent.
Official repository for the AgentVerse framework introduced and evaluated in the paper.
AgentWard prototype repository implementing a lifecycle security architecture for autonomous AI agents with native adaptation to OpenClaw.
Agile V skills repository v1.3, described as the composable AI-agent skills library that operationalizes Agile V and includes context-engineering patterns.
Repository released by the authors for the robust-transformers implementation used with AGRO.
A GitHub CLI extension that exposes a compact interface for reading review threads and submitting inline pull request comments.
Agyn platform for configuring and orchestrating multi-agent systems with explicit communication, roles, and dedicated sandboxes.
Repository for the AI Act Evaluation Benchmark, including structured EU AI Act compliance scenarios, QA pairs, documentation, and scripts.
BabyAGI is used as a representative agentic framework showing how LLMs can be embedded in feedback loops to plan, act, adapt, and manage or prioritize subtasks.
Repository for the violent-python dataset of security-oriented Python code and natural-language descriptions.
Repository for CodeBERT, the pre-trained programming-language model used as the fine-tuned open-model baseline.
Repository containing an implementation related to reproducing 'Clustering by fast search and find of density peaks'.
User interface and implementation for customizing the visual framework to empirical studies using AI accuracy, human adherence, and final decision-making accuracy, and for computing reliance components and Q.
A GitHub Action that reviews pull requests by sending changed file contents to ChatGPT/OpenAI and posting review comments back to the PR.
Repository containing the paper's algorithm code and a link to the associated QuantConnect backtest artifact.
Repository released by the authors for the Writing Quality benchmark, WQRM models, data, and associated evaluation or editing code.
GitHub repository linked by the paper for example Jupyter notebooks and materials related to AIDev and AI teammates in SE 3.0.
Repository for the AIRepr / LLM data-science reproducibility experiments and code released by the authors.
Repository stated by the paper as the public code and data release for AirQA.
Repository for the AISysRev web application, an LLM-based tool for title-abstract screening that imports CSV metadata, applies inclusion/exclusion criteria using multiple LLMs, supports manual screening with LLM guidance, and exports results to CSV.
Contains the source code, dataset-construction scripts, dependency documentation, and result-generation pipelines associated with the paper.
Repository containing code and configurations for training and evaluating AlphaDPO.
Repository containing a practical implementation of the online AdaVol recursive volatility-prediction method.
Public repository containing experimental code supporting the TDQN trading results reported in the paper.
Semantic Kernel is described as a modular, plugin-based framework for integrating LLM and agent capabilities into enterprise and software systems.
Swarm is described as a lightweight multi-agent interface framework for experimenting with multi-agent coordination.
BabyAGI is described as an experimental framework for autonomous task planning and iterative execution through a self-improving task loop.
Author-provided Deformable DETR code repository used for the reconstruction-based SSL evaluation.
Author-provided DETR code repository used for the proposed SSL task experiments.
Official SWE-bench experiment repository used by the authors as the source of execution logs, generated patches, and pass/fail statuses for the studied tools.
Code repository for ACEFormer, the attention-based stock forecasting system introduced by the paper.
Repository for the Online-Mind2Web benchmark and associated evaluation artifacts introduced by the paper.
GPT Engineer is described as an AI coding agent that can generate an entire codebase from a prompt and ask clarifying questions.
MetaGPT is described as a multi-agent framework assigning different GPT roles to form a collaborative software entity for complex tasks.
AgentGPT is described as a framework for rapidly configuring and deploying autonomous AI agents.
Multi-GPT is described as an experimental multi-agent system in which expertGPTs collaborate, communicate, and use short- and long-term memory.
Auto-GPT is described as an early example of GPT-4 running fully autonomously by chaining LLM thoughts to achieve user-set goals.
SuperAGI is described as a developer-centric open-source framework for building, managing, and running autonomous AI agents.
BabyAGI is described as a task-driven autonomous AI agent that builds and prioritises tasks toward an overall goal.
A curated repository of reentrancy attacks from which the paper draws 13 sufficiently isolated real-world exploit contracts for evaluation.
The report states that EnglandCovid is a PyTorch Geometric Temporal dataset and cites this GitHub repository as reference [9].
Repository referenced by the paper for more details on the EnglandCovid dataset conversion and temporal neural network analysis.
A curated repository for the paper's agentic-memory survey, organised around the taxonomy introduced in the manuscript and intended to be updated as the field evolves.
Official repository for the paper, containing a CLAIR preference-generation notebook, cached results, documentation, and links to associated data.
Repository for the AndroidWorld environment, task suite, evaluation logic, and associated experiments introduced by the paper.
Official API-Bank code and data directory containing implemented APIs, datasets, database initialization resources, evaluators, simulators, examples, and demo code.
GitHub repository for the DLS-DDPG implementation used or released with the paper.
Repository for AppWorld Engine and AppWorld Benchmark, including the simulated app environment, benchmark tasks, and evaluation infrastructure.
Repository for APTBench, the benchmark for evaluating agentic potential of base LLMs during pre-training.
Repository identified by the paper as the code for the Arabic tool-calling benchmark work.
AutoResearchClaw is cited as an example of the scenario-verticalized or research-oriented pattern and as the only pipeline/stage subagent case.
deer-flow is cited as an example of the scenario-verticalized or research-oriented pattern.
cline is included as a corpus project and cited as an example of tool-based delegation.
docker-agent is used as a representative orchestration-oriented project combining orchestrator-worker structure, hybrid context, MCP-first tooling, and advanced execution controls.
fast-agent is used as an illustrative project showing tool delegation, hybrid context handling, and MCP-first tool registration.
Gemini CLI is included as one of the official or first-party products from major AI companies in the corpus.
deepagents is cited as an example of the scenario-verticalized or research-oriented pattern.
Mistral Vibe is included as one of the official or first-party products from major AI companies in the corpus.
Kimi CLI is included as one of the official or first-party products from major AI companies in the corpus and is named as a balanced CLI representative.
nullclaw is cited as an example of the Enterprise Full-Featured pattern.
codex is included as an official product/public-evidence case in the corpus and appears in the complete project list.
OpenClaw is analyzed as a corpus project and cited as an example of multi-level recursive or balanced CLI-style infrastructure depending on context.
OpenHands is included in the 70-project corpus and used as an example of event-driven, hybrid-context, enterprise-tool, advanced-safety infrastructure.
agentpool is used to illustrate event-driven subagent architecture with hierarchical context, registry tooling, and enterprise safety classification.
Qwen Code is included as one of the official or first-party products from major AI companies in the corpus.
openfang is cited as an example of the Enterprise Full-Featured pattern.
Repository containing three customer-service chatbot variants generated from different single prompts, with methodology, prompts, and source code for the paper's vibe-architecting demonstration.
GitHub repository for the ARDNS-FN-Quantum code, including the core script and interactive analysis or visualization notebook described by the paper.
Repository containing NeuLR and the paper's appendix or supporting resources for benchmark construction and evaluation.
LongBench is used as the main benchmark source for most datasets and for baseline evaluation settings.
GitHub repository for the Cross-Attention-only Time Series transformer introduced in the paper.
Repository containing the code implementation for evaluating language models using the paper's framework and generating visualizations; described as configurable for other HuggingFace models, metrics, task types, domains, and reasoning types.
Repository containing code for ArgEval, the paper's argumentation-based LLM decision-support framework.
Salesforce repository containing the Art_or_Artifice project folder and associated creativity-evaluation resources.
Repository containing Python implementation files and supplementary materials for detecting hallucinations with specialized model divergence.
Repository reported by the authors as containing the code and data for Ask an Expert / BBMHReasoning experiments.
Repository containing code and scripts for the ASPIRE paper, including captioning, GPT-4 prompt use, spurious object detection, diffusion fine-tuning, image generation, and classifier training workflows.
Tensor2Tensor repository containing the code used by the authors to train and evaluate the Transformer models.
Repository linked by the authors as the code for the attention-retention continual-learning framework.
Author-referenced sample dataset of synthetic identification-document images covering five document types.
Author-referenced repository for generating synthetic document images used in the document identification and information extraction experiment.
Repository containing the implementation of Document Augmentation for dense Retrieval (DAR).
Repository identified by the paper as the official source-code location for the Auto-ADMET method.
Repository containing the Auto-FP benchmark resources, including code, datasets, meta-features, and comprehensive experimental results associated with the paper.
Source-code repository for the AutoFlow framework introduced by the paper.
Repository for the Autoformer model and experiments introduced in the paper.
The paper's AutoGen framework repository for building LLM applications via multi-agent conversations.
Repository for Bias Identification Test in Sentiments, including data files and templates for probing sentiment and toxicity models for bias; the paper specifically uses the disability facet.
Repository for the ADAS codebase introduced by the paper, including the Meta Agent Search implementation and experimental framework.
Repository for ECLAIR, described as enterprise-scale AI for workflows, including code and scripts for the paper experiments and hospital workflow demo materials.
Repository containing folders for datasets, models, plots, pretrained models, additional tests, requirements, and a README with benchmark results for AutoML-DC.
Supplementary repository containing datasets, generated manuscripts, run records, coding runs, and supplemental appendix materials used to support the paper's evaluation.
Code implementation of the data-to-paper framework for backward-traceable AI-driven scientific research.
Hosts the working paper, data.json, figures, tables, and an interactive visualization of Top-N portfolio performance and alpha concentration.
Repository for the BacktestBench benchmark and AutoBacktest implementation, including project folders for AutoBacktest, datasets, figures, tables, environment setup, database setup, and reproduction scripts.
Repository for the BadAgent attack on LLM agents.
Repository for the BagStacking implementation provided by the authors within a Scikit-learn API framework.
AgentGPT is analysed as a general-purpose autonomous LLM-powered multi-agent system with user-guided alignment in selected aspects such as decomposition, agent generation, and resource utilization.
Auto-GPT is analysed as a general-purpose autonomous LLM-powered multi-agent system with autonomous goal decomposition, task action management, and resource utilization.
SuperAGI is analysed as a general-purpose autonomous LLM-powered multi-agent system with some user-guided alignment options for agent-related and resource-related aspects.
BabyAGI is analysed as a general-purpose autonomous LLM-powered multi-agent system with a profile similar to Auto-GPT across many assessed aspects.
The fairseq BART directory provides the implementation interface, released BART checkpoints, and task-specific usage and fine-tuning instructions associated with the paper.
Repository for General AgentBench, the unified benchmark and evaluation framework for general LLM agents.
Official repository for the WorfBench benchmark, WorfEval evaluation implementation, data, and experiment code.
Repository containing code for evaluating MT-BaxBench and MT-SECCODEPLT, the two splits of the MT-Sec evaluation kit.
OpenAI Codex CLI, cited as a lightweight coding agent running in a terminal and evaluated as one of the agent scaffolds.
GitHub repository for the FaithJudge leaderboard and associated faithfulness evaluation resources.
Repository containing the modular benchmarking framework, configurations, scripts, results, and visualization code for comparing retrieval strategies in biomedical RAG.
Repository containing the agent scaffold, behaviour-analysis pipeline, SFT recipes, scripts, utilities, and evaluation suite associated with Behavior Priming for agentic search.
Repository released by the authors with TensorFlow code for BERT, pretrained BERT_BASE and BERT_LARGE checkpoints, and scripts for replicating major fine-tuning experiments.
Official implementation and pretrained sample model for the paper's learned reference-free summary reward, including metric comparison and reward-training scripts.
Official LEAP implementation containing the Python package, prompts, configuration files, data-processing scripts, SFT and DPO training scripts, and evaluation workflows for ALFWorld and WebShop, with InterCode-related code also included in the repository structure.
Repository by genai-analytics containing a beyond-black-box-benchmarking folder with benchmark, core, examples, and sdk subfolders.
ParlAI crowdsourcing code associated with the paper's multi-session model-chat human evaluation.
ParlAI implementation for loading and evaluating the Multi-Session Chat and PersonaSummary tasks introduced by the paper.
Official codebase containing 1dCA data generation, evaluated model implementations, ACT variants, training scripts, evaluation pipelines, and experiment configurations.
Code repository for Multi-Objective Direct Preference Optimization and its experimental workflow.
Repository backing the Agentic Factor Investing project homepage, containing a README, project-framework image, interactive site, and chart data; it is a showcase rather than a complete research-code release.
GoodAI baseline LTM system using a vector database and JSON scratchpad to augment an LLM controller.
Versioned branch of the GoodAI LTM Benchmark repository containing code, test definitions, experiments, result data, and reports corresponding to the paper.
Code for f-DPO with reverse KL, forward KL, Jensen-Shannon, and alpha-divergence regularization, including scripts for IMDB, Anthropic HH, MT-Bench, PPO comparisons, and calibration experiments.
Repository containing the paper's experiments and results for the Agent Assessment Framework.
Repository containing BIG-bench task definitions, JSON and programmatic evaluation infrastructure, documentation, model score files, BIG-bench Lite resources, and contribution workflows.
Repository for the paper's supplemental materials, including the coding schema, all repository metadata, raw graph data, and the LLM coding prompt.
TokenScope is used to extract probabilities for the first decision token where the judge commits to A or B.
Repository containing the MuseD evaluation code and associated research artifact.
Official PyTorch implementation of BOSS for simulated ALFRED experiments associated with the paper.
Project repository for the paper's automated peer-review vulnerability evaluation and adversarial robustness study.
Implementation used to generate the saliency-map explanations that formed the H2 baseline condition.
The gym-sokoban implementation that the authors modified into Sokoban-switch and Sokoban-cells variants for precondition and cost-explanation experiments.
An aggregated dataset of chess opening names and move sequences used by the paper to create opening-position concept datasets.
Source-code repository for the BCDA study and its algorithmic experiments.
OpenDev, the open-source command-line AI coding agent whose architecture, harness, context engineering, tool system, and lessons learned are described in the paper.
Open-source Python package calibrated-explanations, including code repository, examples, notebooks, and evaluation/regression scripts for reproducing experiments.
Source code accompanying the paper, with scripts for decoding, extracting internal-consistency information, and evaluating self-consistency variants.
The open-source CAMEL library introduced by the paper, including agent implementations, prompts, data-generation pipelines, analysis tools, examples, and links to generated datasets.
Repository designated by the paper for code, training conditions, and experimental run records.
JSON benchmark file for the paper's freelance-style task suite.
Repository for the Econometrics-Agent/MetricsAI system that automates econometric analysis through an AI-agent workflow and domain-specific econometric tools.
Official repository directory containing the paper's code, data, and experimental setup for writing alignment through edits.
Repository accompanying the paper, containing the ageval evaluator package, agent tools, baseline evaluators, datasets/labels, experiment configurations, notebooks, and a Gradio app for inspecting annotator outputs.
Official repository containing 50 author-specific writing prompts and anonymized JSON evaluation data organized by quality versus style, prompting versus fine-tuning, and expert versus lay judges.
Repository linked by the paper as the available code for evaluating GPT models on mock CFA exams.
Official implementation of TextGym, evaluated language-agent configurations, EXE, and the benchmark experiments.
Repository containing resources and code for implementing and experimenting with LLMs for vehicle routing problems, including context materials, LLM framework code, oracle algorithms, and verifier code.
Repository for RWE-bench, the benchmark and evaluation environment introduced by the paper for testing LLM agents on real-world evidence generation from medical databases.
Repository containing the FINSABER framework, backtesting code, strategy interfaces, experiment scripts, documentation, and dataset links for reproducing or extending the paper's benchmark.
Official repository containing code, prompts, modules, tests, and post-processing pipeline for generating and evaluating self-generated counterfactual explanations across the paper's datasets and models.
Public repository containing the code, method files, data directory, processing notebook, and framework assets used for the paper's MADR experiments.
LangChain repository audited for the default path from model-produced actions to tool execution.
LangGraph repository audited for mandatory value authorization before ToolNode invocation.
Reference implementation of the paper's deterministic fail-closed authorization gate, including framework integrations and proof scripts.
LlamaIndex repository audited for central dispatch, schema validation, and per-call authorization behavior.
Stripe Agent Toolkit repository audited for client-side authorization of model-supplied payment arguments.
Repository for CAR-bench, including benchmark implementation, tools, task and evaluation workflow, results analysis, and documentation for evaluating multi-turn tool-using LLM agents under uncertainty.
Repository for the paper's cascaded LLM experiments, including implementation details for the framework evaluated in the paper.
BitTern is presented as an open toolkit for low-cost, high-accuracy post-training ternary quantization and 1.58-bit models.
Repository containing the implementation for causal inference via style transfer for OOD generalisation.
Microsoft repository containing the Python implementation and supporting materials for the plug-and-play CoNLI hallucination detection and reduction framework.
Repository for the SVAMP math word-problem benchmark.
Repository for the ASDiv diverse math word-problem dataset.
Repository for the AQuA algebraic word-problem dataset.
Repository for BIG-bench, including the Date Understanding and Sports Understanding tasks.
BIG-bench task repository for the StrategyQA question-only evaluation setting.
Repository associated with the CommonsenseQA benchmark.
Repository for the GSM8K grade-school math word-problem benchmark.
Repository for the Chat2Workflow benchmark and associated workflow-generation/evaluation resources.
Public GitHub location for ChatCollab code and data used or produced by the paper.
MetaGPT repository, representing a prior meta-programming multi-agent framework compared with ChatCollab.
ChatDev repository, representing a prior communicative-agent software-development system compared with ChatCollab.
Repository for the ChatDev framework and the code/data artifact associated with the paper.
Official implementation and configuration repository for the ChatEval multi-agent referee framework and its evaluation experiments.
GitHub repository identified by the paper as the location where the dataset can be accessed.
The paper states that the MMF-Trans code has been open sourced at this GitHub URL, with data requiring authorized access. The repository page itself returned 404 during extraction, so its contents could not be inspected.
ClawNet repository containing the governed multi-agent social network implementation, including core/gateway, server, desktop, macOS client, and setup components.
Repository for CMAT: A Multi-Agent Collaboration Tuning Framework for Enhancing Small Language Models, including code, configs, data, docs, and README material describing the framework and evaluation setup.
A GitHub repository listed by the paper as an accompanying curated resource for papers on code as agent harness.
Repository containing released benchmark data derived from SWE-CARE, pipeline scripts for filtering, environment building, test generation, agent resolution, and tool evaluation, plus compressed raw experimental outputs.
Repository for the CodeAssistBench benchmark, datasets, prompts, scripts, and evaluation framework for AI coding assistants on real GitHub issues.
Open-source repository associated with the CodeScout model family and RL recipe for code localization agents.
GitHub repository for CodeTaste, including benchmark infrastructure, agent/evaluation scripts, documentation, and links to benchmark artifacts and precomputed outputs.
Repository for Codev-Bench, the developer-centric repository-level code-completion benchmark constructed with Codev-Agent.
Repository titled KotiJaddu/Masters-Project, described on GitHub as code supporting the author's Master's thesis, with Python source and model folders.
Repository for the LLMWorkflowGenerator project, containing the Python Controller implementation, Android/Termux-oriented workflow automation code, sample files, experiment outputs, and setup instructions.
Repository provided by the authors for reproducing the LLM, human-comparison, and algorithmic bandit experiments.
Repository for the Bitcoin Fee Rate Prediction Project, including data, model notebooks, scripts, result outputs, plots, and citation information for arXiv:2502.01029.
Repository reported by the paper as the code for confidence estimation in LLM-based dialogue state tracking.
Official repository and Python package for CONFLARE, supporting document loading, cleaning, chunking, calibration-set creation or loading, and conformal retrieval-augmented generation.
Repository linked by the authors as the public code for experiments in Conformal Prediction as Bayesian Quadrature.
Repository linked by the paper for CCA/SWE-Bench-related materials.
Open-source PyTorch repository from which the authors selected reproducible GitHub issues requiring specialist debugging.
Python repository for generating PortBench portfolio-theory tasks and evaluating LLM portfolio decisions.
Repository accompanying the paper with selected Python code for feature construction, metrics, and the model. The README states that the full proprietary backtesting and continuous-futures processing framework and licensed data are not included.
Repository released by the authors for constructing hierarchical NAS spaces based on CFGs and reproducing the BOHNAS experiments.
GitHub repository identified by the authors as containing the standard performance forecaster model for the Allora network.
Repository directory containing the implementation and resources for the CoT-MAE contextual masked auto-encoder introduced and evaluated in the paper.
Repository containing the paper's data and codebase, including data collection materials for the user feedback and annotation tasks.
Python repository containing slow-tail, V-shape, and combined max-cash modules, execution scripts, data-schema documentation, and reproducible output paths for the empirical cash-overlay study.
Repository containing code, data artifacts, and reproducible empirical materials for the continuous smooth-signal growth-versus-defensive allocation framework.
Repository identified by the paper as the posted model code for reproducing the ConFIRM workflow.
GitHub repository associated with the paper's cooperative knowledge distillation method.
Official CooperBench repository containing the benchmark package, dataset tooling, task-running CLI, evaluation logic, and cooperative/solo/team settings for coding-agent experiments.
Repository containing the implementation and configuration details for LSNPC experiments.
Repository containing the implementation of CryptoMamba, baseline models, data preprocessing, model training, evaluation metrics, and trading simulation scripts.
RAGQALeaderboard is used as a benchmark environment for evaluating RAG-QA performance across multi-hop, single-hop, and biomedical question-answering tasks.
Source-code repository for DAG-MoE, including the learned structural aggregation module evaluated in the paper.
IBM Agentics is the agentic AI framework on which DAO-AI is built; the paper uses its ATypes, logical transduction, and scalable workflow concepts.
GitHub repository identified by the paper as containing DAO-GP source code and datasets.
FastChat v0.2.5 question JSONL containing 80 diverse evaluation queries.
Repository implementing data-local, ensemble-based, LLM-guided NAS for multiclass multimodal time-series classification, including local executor scripts, remote LLM proposer scripts, schemas, and a Flask control interface.
Repository cited by the paper for NVIDIA<ef><bf><bd>s Comprehensive Verilog Design Problems benchmark, which provides the selected code-generation and code-comprehension tasks used to evaluate SLMs and LLMs.
Repository for the DB2-TransF time-series forecasting model introduced and evaluated in the paper.
Code repository for the limited teacher supervision decoding method introduced in the paper.
AllenNLP contains the official PyTorch ELMo module and scalar-mix implementation for computing the paper's contextual representations.
Official TensorFlow implementation for training the pretrained bidirectional language model and computing ELMo representations introduced by the paper.
Repository containing the data-pipeline code, hedging simulator, actor-critic training workflow, saved configurations and models, notebooks, evaluation utilities, tests, figures, and paper materials for reinforcement-learning-based hedging of SPX and SPY exposures.
Public repository containing software code and datasets or scripts for reproducing the baseline algorithms and experiments.
Repository stated by the paper to contain the market simulation and experiment code.
Repository linked by the paper as containing the publicly available datasets and code used for the DRL-in-finance experiments.
Repository containing code associated with the paper's A2C, DDPG, PPO, ensemble-selection, and backtesting workflow.
Repository containing source code, datasets, and supplementary materials associated with the multi-horizon NEM electricity price forecasting benchmark.
The repository contains DeepAries source code, market data folders, checkpoints, model components, experiment code, and instructions for training and inference.
Repository path containing DeepPlanning benchmark code, travel and shopping domain runners, evaluation utilities, configuration files, and instructions for reproducing benchmark results.
The DeepSpeed repository contains the framework, code, tutorials, and documentation for large-scale model training and inference, including DeepSpeed-MoE components.
Repository for DeepTraderX, a deep-learning trading agent running in Threaded-BSE.
Threaded Bristol Stock Exchange repository providing the asynchronous market simulator and working trading agents used for DTX training data and experiments.
Repository identified by the paper as containing code and data for the finance hallucination experiments.
Code, documentation, and demos for ToolUniverse, the ecosystem for building AI scientists from language models, reasoning models, and agents.
Official implementation repository released for the Demonstrate-Search-Predict framework.
Salesforce AI Research repository for FinDAP: Demystifying Domain-adaptive Post-training for Financial LLMs, including framework materials, training scripts, evaluation guidance, and links to model/data artifacts.
Repository associated with the AgentFail dataset and website for failure lifecycle data and analyses.
Repository containing DPR training, retrieval, evaluation, data-processing tools, configurations, and released model resources.
GitHub repository for the TSCC2019 competition data used as real Hangzhou traffic data in the paper's traffic-light-control benchmark.
Repository containing the blank Design-OS template, design-case artifacts, prompt files, simulation code, plots, and verification report used to support reuse and replicability testing.
GitHub repository for the TalkTuner paper, including code and data for a dashboard that visualizes and controls a chatbot LLM's internal user model.
GitHub repository for the evolutionary multi-objective neural architecture search approach introduced in the paper.
Repository linked by the paper for reproducing or inspecting the OAS generation experiment.
Repository describing DeXposure-FM as a time-series graph foundation model for forecasting inter-protocol credit exposure, with scripts for experiments, macroprudential tools, checkpoints, and dataset helpers.
The code URL reported in the paper's v2 HTML/PDF abstract for the DeXposure-FM project.
Code repository for the DiffLOB regime-conditioned diffusion model.
Training code for supervised fine-tuning followed by DPO preference learning on causal Hugging Face language models, with dataset and trainer utilities.
Google Java Format repository; the paper uses Newlines.java from this repository for generated-vs-original unit-test comparison.
Junit5 Modular World sample module; the paper uses a Flavor.java code snippet from this repository for prompt-based test generation and comparison.
Contains code, data, configurations, and scripts for the FineLogic training and evaluation experiments.
Author repository for the GSM8K-AI-SubQ reasoning dataset and baselines for distilling LLM decomposition abilities into compact language models.
Official source code repository for the Distilling step-by-step method introduced in the paper.
Repository containing code for the supervised and unsupervised LLM uncertainty experiments reported in the paper.
GitHub repository released by the authors for Distributed Conformal Prediction via Message Passing.
Official code repository for the Division-of-Thoughts framework introduced and evaluated in the paper.
Official self-contained SayCan implementation in a simulated tabletop environment using a UR5 setup, ViLD affordances, GPT-3 planning, and a CLIPort pick-and-place policy.
Official SayCan dataset files mapping natural-language user instructions and initial conditions to possible solution plans.
Repository containing code and notebooks associated with the S&P 500 graph-neural-network forecasting project.
Repository containing code and data for evaluating uncertainty estimation in LLM instruction-following.
Repository reported by the paper as containing the prediction model implementation and experiment data.
Repository for the DocAgent multi-agent code documentation generation framework introduced by the paper.
Repository for the code, human-subject study materials, results, and supplementary materials associated with the Persona framework and AAAI 2025 paper.
Semantic routing package for routing inputs by embedding or intent similarity.
Framework using repeated generations, verification prompts, and confidence estimates to decide whether to escalate to larger models.
AWS multi-agent orchestration framework that includes prompt-based routing or agent selection patterns.
Implementation associated with routing prompts to pre-trained experts after fine-tuned meta-model categorisation.
Implementation associated with deciding whether a query requires a complex prompting strategy.
Iterative multi-agent code generation system using execution success as a routing signal.
LLM routing implementation associated with assessing model adequacy through multiple responses and ground-truth comparison.
Routing-agent implementation using synthetic data and small classifiers for classification-based routing.
Orchestrator implementation using decoder-only LLM representations for routing or model selection.
Framework for serving and evaluating routers that choose between LLMs using preference-oriented routing strategies.
Task-planning framework in which an LLM selects among models or tools based on descriptions and user tasks.
Implementation assessing consistency across reasoning representations for cascade-style routing.
OpenAI multi-agent orchestration framework discussed as an example of prompt-based routing practice.
Fine-tuned model framework for API call generation, discussed as treating routing as a code generation problem.
Framework for reducing LLM application cost using LLM cascades and related strategies.
Adaptive RAG framework that routes among no retrieval, single-step retrieval, and multi-step retrieval paths according to query complexity.
Code and data for a multi-LLM routing benchmark and evaluation framework.
GitHub repository containing the Indonesian financial-domain language-model code and post-trained IndoBERT models.
Implementations of original, sparse, continuous, and related statistical jump models.
Code and model artifacts for Reinforced Token Optimization, the paper's DPO-derived token-reward and PPO/RL alignment method.
Official repository containing the DS-1000 data, execution-based evaluators, environment files, inference scripts, and released baseline results.
Repository for evaluating LLMs on the DSBC dataset, including response generation, LLM-as-judge evaluation, command-line usage, and dataset evaluation utilities.
Code repository for the DSTCGCN traffic forecasting model introduced and evaluated in the paper.
A scikit-learn-style implementation of a collection of statistical jump models, including the model family used to identify asset-specific market regimes.
Repository indicated by the paper as the code website for DVGNN.
Repository containing experimental analysis, tools, and code associated with dynamic design of machine-learning pipelines via metalearning.
The first author's repository providing implementations of statistical jump models, including methods relevant to the sparse jump-model analysis used in the paper.
Official GitHub repository containing code for the paper Dynamic Graph Convolutional Network with Attention Fusion for Traffic Flow Prediction.
Repository for the Dynamic Meta-Learning for Adaptive XGBoost-Neural Ensembles implementation, including source code and test datasets as described by the paper.
Python repository containing pseudo-answer generation, silver-label construction, input processing, multi-task training, inference, and evaluation scripts for the six reported datasets.
Repository containing the EasyRAG pipeline, ingestion code, retrievers, rerankers, prompt templates, challenge scripts, Docker deployment, FastAPI service, Streamlit WebUI, and processed challenge assets.
Official Microsoft Research implementation of the Sui Generis scoring pipeline introduced in the paper.
Public GitHub repository for the study's replication materials, mirrored in Zenodo.
Code repository for CAID, including the multi-agent workflow where a central manager delegates tasks to engineer agents that run asynchronously in isolated git worktrees, plus scripts and task modules for Commit0 and PaperBench experiments.
NovGrid extends MiniGrid with a generalized novelty generator so environment properties and dynamics can change and agents can be evaluated on adaptation to those changes.
OWL is compared against Efficient Agents in the agent framework evaluation.
Smolagents is compared against Efficient Agents and OWL in the agent framework evaluation.
Repository linked by the paper as its code resource, associated with the OAgents/Efficient Agents work.
Public code repository for the DASH architecture-search algorithm introduced and evaluated in the paper.
fairseq example code and resources for training and evaluating the paper's MoE language models.
Source-code repository for Helium, the workflow-aware LLM serving system proposed and evaluated in the paper.
Repository for the Temporal Neural Common Neighbor model introduced in the paper.
Official repository containing scripts and links to reconstruct the ELI5 dataset, build support documents and splits, format multi-task data, train and evaluate the reported models, and use the released pretrained checkpoint.
Repository for EmbedLLM materials, described by the paper as containing the dataset, code, and embedder for further research and application.
Repository for the Embodied Web Agents project, including web environment hosting instructions and model-running folders for indoor, outdoor, and geolocation tasks.
Python SDK that lets users obtain drone observations and issue control actions through the Embodied City online API.
Repository containing simulator-related materials, datasets, task code, prompts, VLN code, and documentation for the EmbodiedCity benchmark.
Code and experiment resources for measuring how context-characteristic sensitivity changes across instruction-fine-tuning stages.
Repository named Emoji-Embedding-For-Finance, cited throughout the paper as the source for model-vs-BERT comparisons, emoji frequencies, BTC/VCRIX figures, sentiment time series, and trading-strategy outputs.
Official repository for the paper, containing code and data artifacts for generating analysis-report features and training the hybrid asset pricing model.
Official code and data repository for the ALCE benchmark and evaluation framework introduced in the paper.
Pre-release Python/PyTorch code for building chart-image datasets, splitting data, training the ResNet trader, and inferring triple-I weights.
Repository for implementing Iter-CoT, the paper's iterative bootstrapping method for chain-of-thought prompting.
Preferred Multi-turn Benchmark for Finance in Japanese, used by the paper to evaluate generation quality across financial dialogue tasks.
Pathway is used as the vector store implementation in the paper's retrieval system.
Stated repository for the modular Python prototype implementing the neuro-symbolic ontology-based LLM validation pipeline.
Repository containing implementation for the PSX interpretability algorithms and models discussed in the paper.
Repository indicated by the paper as containing code for AdvDistill; the URL returned 404 when fetched during extraction.
Public implementation of the structured-matrix Transformer enhancement framework introduced in the paper.
Repository containing the LoT prompting implementation, scripts for CoT and LoT experiments, requirements, and quick-run instructions.
GitHub repository cited by the paper as the source of the second obesity classification dataset.
Repository reported by the paper as containing code, curated data pointers, generated figures, and result tables for the reproduction and robustness analysis.
Implementation of faithfulness-aware decoding used for the advanced-decoding baseline.
FactPEGASUS code and models used to test whether the proposed augmentation transfers to another factuality-aware contrastive pipeline.
Official CLIFF implementation used as the contrastive-learning baseline and as the training framework combined with the proposed counterfactual augmentation.
Repository hosting the ESG-FTSE corpus of news articles with ESG relevance labels.
Repository released by the authors containing code for the agentic benchmark assessment and related experiments.
Repository associated with the paper's contamination-detection work for LLM evaluation. The paper links it as code and data; the currently visible README describes a lightweight tool for identifying and analysing potential contamination without access to LLM training data.
Repository URL listed by the paper for the implementation of the confidence IQN experiments.
A production-grade evaluation toolkit for LLM agent outputs, including metrics for diversity, reliability, cascade uncertainty, perturbation consistency, consistency, factual grounding, hallucination, explainability, and drift.
Repository containing the AGENTbench harness used to evaluate coding agents under NONE, LLM, and HUMAN repository-level context settings on AGENTbench and SWE-bench Lite.
Public GitHub repository released by the authors to support reproducibility and further research on financial relationship graph evaluation.
Repository described as a resource hub for understanding, detecting, and mitigating biases in financial-domain LLMs, including a Structural Validity Checklist, a literature review dashboard, and an automatic bias detection dashboard.
Repository associated with the M4 competition dataset and methods, used by the paper for benchmark data and base learner forecasts.
Repository containing the source code for the paper's FFORMA and ES-RNN ensemble experiments.
Official repository containing the LoCoMo data release, conversation-generation code, prompts, and evaluation scripts.
Phoenix is cited among tools that provide analytics and evaluation orchestration capabilities for agent or LLM evaluation.
DeepEval is listed among tools that support analytics, evaluation orchestration, and debugging for LLM or agent evaluation workflows.
OpenAI Evals is discussed as an open-source framework for specifying evaluation tasks and metrics and automating execution and reporting.
HAL is cited as a holistic agent leaderboard or harness for centralized and reproducible agent evaluation.
Inspect AI is cited as a framework for large language model evaluations and included among evaluation tooling examples.
Repository named in the paper for ICD coding explainability evaluation resources, including RD-IV-10-related artifacts and generated rationales.
Paper-provided implementation link for the evaluation-awareness scaling-law experiments; the arXiv HTML resolves through an Anonymous Github mirror.
R code, datasets, and experiment scripts for EvoAAA, the evolutionary autoencoder architecture search methodology introduced in the paper.
Repository containing code and technical details for the Multi-Agent Scoring System for essay assessment.
Repository providing BanditBench and inference code; the paper also notes installation via pip install banditbench.
Official repository for the CodeAct framework, evaluation code, CodeActAgent deployment components, scripts, and links to the released data and models.
Repository identified by the paper as containing the code and data for the software-developing agent framework evaluated in the study.
The repository is linked by the paper as the location for data and code and contains files including a notebook, UNSW-NB15 dataset archive, and feature metadata.
Repository named by the authors as containing all code for the experiments that apply feature importance, SHAP, and LIME to PPO portfolio-management predictions.
Existing Reinforcement learning in portfolio management repository that the authors identify as the starting framework for integrating explainability into a PPO-based model.
The cpath package implements the paper's counterfactual-path method for R and Python.
UnifiedSKG is the structured knowledge grounding model used in the paper's text-to-SQL case study.
Repository containing the benchmark data, task files, representative logs, and evaluation scripts for AutoGen, MetaGPT, and TaskWeaver.
GPT-Engineer is listed as a code-domain LLM-based single-agent system and cited as a GitHub project that generates code repositories from prompts.
GPTresearcher is listed as a research-domain LLM-based autonomous agent for online comprehensive research.
AIlegion is listed as a universal LLM-powered autonomous agent platform.
LoopGPT is listed as a universal modular Auto-GPT-style LLM agent framework.
AGiXT is listed as a universal AI automation platform with instruction management, memory, and plugins.
LangChain is discussed as an open-source framework supporting LLM-based agent software development and tool integration.
DemoGPT is listed as a code-support LLM-based agent system for creating LangChain applications through prompts.
AgentGPT is discussed as an agent framework offering browser-based assembly, configuration, deployment, fine-tuning, and local data incorporation.
Auto-GPT is discussed as an open-source agent template/framework for decomposing objectives and executing tasks in a loop.
SmolModels / smol-ai developer is listed as a code-domain LLM-based agent system with self-feedback and tool use.
WorkGPT is discussed as an open-source GPT agent framework for invoking APIs.
SuperAGI is listed as a universal open-source autonomous AI agent framework.
XLang is discussed as an open-source framework for building and evaluating language model agents through executable language grounding.
BabyAGI is listed as an LLM-based agent that creates tasks from objectives and stores or retrieves task results.
BabyAGI is described as an OpenAI-powered task management system that uses vector databases such as Chroma or Weaviate to manage, prioritize, execute, store, and recall task-related information.
Code for defining text-to-text tasks and mixtures, preprocessing and evaluating datasets, training and fine-tuning T5 models, and reproducing the paper's experiments, with links to released checkpoints.
Repository for ExpNote: Black-box Large Language Models are Better Task Solvers with Experience Notebook, including experiment code, datasets, scripts, and setup instructions.
Repository containing the 50-claim development dataset, the 25-claim post-finalization dataset, and dataset statistics used in the paper.
Repository containing materials from Stages 1, 2, and 3 of the expert formalization study evaluating the super-pattern.
Repository containing the paper's released Self-Contrast code and data.
Open-source implementation of the paper's GPT-2 generation, candidate ranking, and training-data memorization attack workflow.
Repository containing FactReview, RefCopilot, demos, source code, scripts, tests, and documentation for evidence-grounded reviews of ML papers.
Repository containing pilot and main JSON datasets, prompt generation code, evaluation code, prompt templates, and unit-group definitions for the FAITH paper.
PyTorch code, configurations, training and unmasking scripts, evaluation utilities, and pretrained checkpoints for ImageNet 256x256 and 512x512 MaskDiT models.
Code repository linked by the paper for the FedSPM method and experiments.
Official repository for reproducing the paper's output-refinement and policy-refinement experiments.
Repository associated with the Cross-Attentive Time-Series Trend Network described and evaluated in the paper.
Open-source repository associated with PIXIU and FinBen, containing financial LLM resources, evaluation datasets, benchmark materials, code, and links to related models and leaderboards.
Repository indicated by the authors for the FinBERT2 work, including the specialized encoder and related downstream variants or resources.
ProgramFC repository used as the implementation basis for the LLM-based composite fact-checking system in the experiment.
Official repository for offline reward-model training and preference-based language-model fine-tuning, including smaller fine-tuned models and a subset of collected human labels.
GitHub repository for a Chinese financial news sentiment classification dataset containing train and test CSV files and financial news sentiment labels.
Repository containing code for FinEAS and the paper's BERT, BiLSTM, and FinBERT financial-news sentiment experiments, together with reported result tables.
Repository containing data preprocessing code, FineFT algorithm training/validation/testing scripts, VAE routing components, trading-environment implementation, baseline code, and analysis utilities.
Python source code repository for FINMEM, the LLM trading agent with layered memory and character design.
Repository for the FinReport code and datasets released by the paper.
Astock dataset used for stock and news data, train-validation-test splitting, OOD analysis, and backtesting.
Repository for the paper's dataset construction pipeline, data collection module, FinRpt framework modules, benchmark evaluation code, fine-tuning setup, reinforcement-learning setup, and website front-end code.
GitHub repository associated with the paper's benchmark for online financial RAG evaluation.
Repository released by the authors for FinTexTS framework code and pilot study implementation.
Open-source implementation repository for FinWorld, the end-to-end financial AI research and deployment platform introduced in the paper.
Contains FireAct prompts, task and tool definitions, trajectory-generation and evaluation code, fine-tuning scripts, example training data, and model references.
Repository containing the Flow multi-agent workflow automation implementation, including workflow management code, prompts, validators, notebooks, and generated examples.
Code repository for the FlowAgent framework introduced by the paper.
Official repository containing FlowBench source data organization, turn-level and session-level evaluation code, scripts, prompts, and setup instructions.
GitHub Gist with a replenishment plan generated by the Flowr DC Replenishment Planning Agent, including outlet allocations, route assignments, vehicle assignments, route summary, consolidation opportunities, and human review queue.
GitHub Gist with a purchase-order report generated by the Flowr Procurement and Ordering Agent, including order quantities, supplier justifications, delivery estimates, consolidated supplier orders, and human review flags.
Repository made available by the authors for the proposed models used to forecast extreme Bitcoin volatility movements.
Repository for the FiGASR package and related sentiment-indicator resources.
Repository containing Python code for fine-grained aspect-based sentiment analysis in the economic news setting.
Public multivariate time-series dataset repository containing the exchange-rate data used as real-data-E.
Lag-Llama pretrained model repository used as the probabilistic zero-shot forecaster in the experiments.
Official implementation of the Diffusion Model-Based Predictor baseline for robust offline RL against state-observation perturbations.
Repository containing the implementation and results for LSTM and GRU time-series forecasting experiments discussed by the paper.
Repository containing FoT code, task and benchmark scripts, datasets, model and method modules, and evaluation instructions for Game of 24, GSM8K, MATH-500, and AIME.
Repository containing more detailed prompts and implementation code for the LAWN antenna self-evolution framework discussed in the paper.
Repository stated by the authors to contain the complete source code employed in the research.
Repository for the paper 'From Coarse to Fine: Self-Adaptive Hierarchical Planning for LLM Agents', containing implementation-related folders such as agents, configs, envs, plan generation, prompts, tasks, utilities, and experiment scripts.
Repository for the paper From Emergence to Control: Probing and Modulating Self-Reflection in Language Models, describing probing vectors, model insertion, and reflection analysis for controlling self-reflection in LLMs.
A maintained paper list and resource collection about LLM-as-a-judge.
Repository for the paper's semi-automated ontology and knowledge-graph construction pipeline, including prompts, code, data, generated artifacts, results, and evaluation materials.
Repository containing the reproducibility-related publication data and analysis code from the authors' prior biodiversity deep-learning study, which supplied the source dataset for the current pipeline.
Repository titled for the paper 'From Large AI Models to Agentic AI: A Tutorial on Future Intelligent Communications' and identified by GitHub as the code repository for the paper.
BeeAI is described as the experimental platform central to IBM's ACP, supporting local-first orchestration, agent discovery, REST endpoints, SDKs, telemetry, and multi-agent execution.
The MCP servers repository is cited as an ecosystem of reference and integration servers for file management, databases, Google Drive, Git, GitHub, GitLab, Slack, Google Maps, image generators, and search APIs.
OpenAI Swarm is reviewed as a lightweight, stateless abstraction for multi-agent systems with agent definitions, dynamic handoffs, context management, direct function calling, streaming, and backend flexibility.
Repository reported by the paper as containing code and data for the proposed news-aware LLM time series forecasting framework.
Repository containing code and data for the HippoRAG family, including the framework introduced and evaluated in this paper.
Repository path for a Workflow Composer that transforms natural-language research questions into executable HyperFlow workflows for the 1000 Genomes Project, including source code, tests, Skills/resources, CLI commands, and evaluation-related materials.
A companion GitHub repository associated with the survey on agentic workflow optimization.
Repository for the paper containing raw experimental data, reasoning processes and outputs, keyword-statistic results, LLM-as-a-judge scores, and score-analysis spreadsheets.
Repository associated with the paper<ef><bf><bd>s recursors/refiners work, containing generated Prolog programs, models, execution traces, and related materials for the implemented system.
AWorld-RL is the linked project repository that houses FunReason-MT materials and describes the introduced multi-turn function-calling data-synthesis framework.
Repository identified by the paper as containing the pre-processed dataset, raw news article data, and implementation code for M2VN.
Repository containing GAAMA's typed memory graph, semantic and PPR retrieval components, prompts, storage adapters, and LoCoMo evaluation scripts.
Repository for the Game-theoretic LLM paper, including source code, setup instructions, complete-information game experiments, workflow experiments, and Deal-or-No-Deal experiments.
Repository reported by the author as containing the code and dataset used in the paper.
ytopt is a machine-learning-based autotuning and hyperparameter-optimization framework used in Section IV-B to search coefficient spaces for the reviewed scaling laws.
The Lila benchmark repository supplies the program-form mathematical reasoning data used in the paper's mathematical program-synthesis experiments.
Public repository for the paper's generative-agent architecture and Smallville simulation code.
PyTorch repository containing generative sparse-index-tracking code, comparison code, a GECCO 2023 directory, backtesting material, and instructions for obtaining the associated dataset.
Official repository containing Gorilla inference resources, APIBench data, evaluation scripts, model outputs, and materials for reproducing the paper's results.
OpenAI Evals, a framework for creating and running model benchmarks and inspecting performance sample by sample.
Public prompts and code for the GPT-based automated review generation workflow used in the study.
Repository linked by the paper as the implementation of Grammar Search for multi-agent systems.
Source code for GraphGPT and the Graph Eulerian Transformer workflow.
RAGChecker is an open-source visual analytics tool that compares LLM outputs against source documents, supports GraphEval+ and SICI-style detection, and presents claim reliability through an interactive quadrant-based visualization.
Repository for the GTA benchmark, dataset, code, and evaluation materials for general tool agents.
Official repository for H3M-SSMoEs containing Python code, model components, training and backtesting scripts, and links to datasets and model weights.
Repository stated as the release location for HDFlow code and data.
Repository associated with the paper's released LOFin benchmark and HiREC implementation.
Repository containing code to reproduce the HCNN experiments reported in the paper.
Repository containing annotated transcripts of each participant made available for future studies.
TheAgentCompany repository contains sandboxed work environments, task directories, evaluators, task instructions, and supporting files for many task instances listed in the paper's appendix task table.
Repository for coding-agent token-consumption analysis, including dataset-building scripts, multi-model analysis scripts, phase-level token decomposition, and self-prediction correlation computation.
Python codebase for running the ensemble generalization-gap experiments, computing dataset complexity metrics, and generating per-dataset outputs and figures.
GitHub directory linked by the paper for the deterministic prediction task source code, including Pauli string multiplication, divide-and-conquer, letter replacement, and addition-related files.
Repository containing the dataset and experimental code for the study of memory addition, deletion, and experience-following behaviour in LLM agents.
The repository provides code folders for HAG-XAI object detection, HAG-XAI image classification, FullGradCAM for Yolo-v5s, and FullGradCAM for Faster-RCNN, along with links to experimental materials, human attention data, and pretrained model files.
Repository containing data and analysis code for the human-alignment AI-assisted decision-making study.
Contains the Auto_Driving_Highway code, prompt materials, results, and instructions for training an RL agent with an LLM in the reward loop.
Azure Verified Modules is used by the paper as a cloud-native curated module ecosystem that demonstrates governance, standards, testing, versioning, and consistent interfaces.
Repository indicated by the paper for prompts, code, anonymised data, and the full questionnaire related to the AI Narrative Test.
Repository for the S&P 500 membership forecasting project, including EDA.ipynb, Modeling_Process.ipynb, model-interpretability figures, and Python dependencies.
Source code for the HybRank hybrid and collaborative passage-reranking model.
Official implementation repository for HyperAgent, including the multi-agent software engineering framework and benchmark reproduction scripts.
GitHub repository for the ICICLE method introduced and evaluated in the paper.
Repository containing ToolEmu code, emulators, evaluators, curated toolkit and test-case assets, scripts, and notebooks.
GitHub repository released by the authors for IL-PCSR and developed models.
Complete codebase and setup instructions for the sentiment-driven stock prediction experiments.
Repository maintained by the authors as an official page and living resource list for the survey on implicit reasoning in LLMs.
Open-source repository cited as the source for the Reuters & Bloomberg news-title stock-prediction dataset used to build headline vectors.
Repository titled Anote-Text-Classification containing dataset folders, trial runs, requirements, README explanations, and model performance comparisons for GPT-3.5 Turbo, SetFit, and BERT across the evaluated datasets.
Repository for the CARE native retrieval-augmented reasoning framework, including scripts, evaluation resources, documentation, and training examples.
Official FactCC implementation and model used to score whether generated summary claims are factually consistent with source documents.
Preliminary implementation of the paper's multiagent debate experiments, with code for arithmetic, GSM8K, biography, and MMLU tasks.
Repository for Fusion-in-Decoder, providing the Natural Questions version augmented with DPR-retrieved passages that RETRO uses for its question-answering experiment and representing a principal comparison system.
NexusRaven-V2 repository linked by the paper for the Nexus function-calling evaluation benchmark.
Repository containing notebooks for SFT and LoRA-based DPO training, generation and reward-model evaluation of the four OPT-350M variants, evaluation data and JSON outputs, plots, and examples of noisy preference pairs.
Code repository for the paper's QE-based machine-translation feedback-training experiments and RAFT+ method.
Code repository for the MAP paper, including implementations and evaluation scripts for Tower of Hanoi, CogEval graph tasks, PlanBench, StrategyQA, and transfer experiments.
Microchain is the agent-based library used to inject function declarations, descriptions, and examples into prompts, call functions during the reasoning loop, and record assistant/user chat-completion chains.
Public repository for InlineCoder, the framework introduced and evaluated in the paper.
Repository linked by the paper for InfoMosaic-Bench/InfoMosaic-Flow resources.
Published ETT dataset collected by the authors and used as a core benchmark dataset in the paper.
Source code for the Informer model and experiments.
Aider source repository analyzed for user-driven loop design, repo-map retrieval, edit formats, and summarization behavior.
OpenCode source repository analyzed for tool interface design, event bus, SQLite session persistence, dynamic tools, and role-based sub-agents.
Moatless Tools source repository analyzed for MCTS orchestration, tree-structured state, action classes, semantic search, in-memory shadow mode, and actor-critic routing.
AutoCodeRover source repository analyzed for phased scaffold control, search-only tools, AST-aware retrieval, SBFL, and Docker-based execution.
Cline source repository analyzed for recursive control flow, IDE coupling, shadow git checkpoints, delegation, and LLM-initiated compaction.
Prometheus source repository analyzed for LangGraph-based phased control, graph-scoped state, knowledge graph retrieval, per-node tool scoping, and multi-tier persistence.
Gemini CLI source repository analyzed as one of the 13 coding agent scaffolds.
Codex CLI source repository analyzed for event-driven ReAct control, dynamic tool rebuilding, sandboxing, Guardian safety routing, memory extraction, and sub-agent delegation.
Agentless source repository analyzed for fixed-pipeline architecture, hierarchical localization, sampling, and JSONL pipeline state.
OpenHands source repository analyzed for event-sourced architecture, tool interfaces, Docker execution, and delegation mechanisms.
mini-swe-agent source repository analyzed as a deliberately minimal baseline scaffold with a single bash tool and simple ReAct loop.
SWE-agent source repository analyzed for ReAct control loop, tool bundles, retry behavior, Docker isolation, and compaction processors.
DARS-Agent source repository analyzed for depth-first tree search, SWE-agent-derived tools, greedy LLM critic selection, and Docker reset/replay state recovery.
A topic-agnostic, provenance-first pipeline for legitimate AI-assisted drafting of a Related Work section using human-authored structured notes, strict no-invention constraints, interaction logs, generated taxonomy, draft output, audit table, provenance card, and related artifacts.
Repository associated with the paper's survey of instruction-tuning research.
Python repository associated with hybrid physics-based and data-driven building energy modeling, with folders for evaluation, feature generation, forecasting, models, preprocessing, and utilities.
Repository containing code notebooks, data files, tickers, metrics/metadata, SQL queries, and the list of companies associated with the paper's algorithmic stock-market trading system.
Repository for the MiniGrid gridworld environment used for the paper's four-room and six-room navigation experiments.
Repository for the FastAPI control server, callback-augmented interactive trainer, React/TypeScript dashboard, examples, and LLM-based tuning demonstration.
The paper states that the ICM protocol is open source under the MIT license and that referenced workspaces are available or buildable through this repository.
Repository identified by the paper as containing the code for the GDELT headline extraction, FinBERT sentiment scoring, feature engineering, modeling, and backtesting workflow.
Public code repository for the paper's context-faithfulness experiments.
Repository for InvestLM, the LLaMA-based financial-domain instruction-tuned model released to the research community under the same licensing terms as LLaMA.
Python project for forecasting variance-covariance matrices in a Markowitz framework across cryptocurrency and traditional asset markets; the repository README states that it produced the paper's results.
Code and data repository for reproducing the paper's evaluator training, SFT+ILR, SFT+DPO, naive-ILR, and supervision-quality experiments.
Repository containing Stage 1 JEPA+DAAM training, Stage 2 decoder training, FSQ and mixed-radix packing, a HiFi-GAN decoder with optional DAAM gating, and DeepSpeed integration.
Repository associated with the Jr. AI Scientist system; the paper points readers there for issues, comments, questions, and planned codebase release.
Official K-COMP codebase with retrieval, data-processing, training, inference scripts, environment configuration, short medical descriptions, and links to model checkpoints and datasets.
Repository path containing code or supplementary material for the shell-based keyword-search agent used in the paper.
Official implementation of the KnowAgent framework, including HotpotQA and ALFWorld path-generation scripts, trajectory filtering and merging, and LoRA-based knowledgeable self-learning.
Repository containing the question-generation compression workflow, paper-card generation, syntactic multihop experiments, datasets, evaluation scripts, and reported result files.
Official repository released with the paper, containing extensible implementations of KTO, DPO, offline PPO, ORPO, and other human-aware loss functions.
Repository for the technical report, providing code for high-frequency trading prediction under label imbalance, including MLP, LSTM, BERT, and Mamba backbones plus class weighting and data balancing options. The original dataset is not provided because of copyright restrictions.
Official LaMMA-P repository containing code, PDDL resources, scripts, Fast Downward submodule setup, and MAT-THOR test dataset files for running the method in AI2-THOR.
Official repository accompanying the GPT-3 paper, containing synthetic arithmetic and word-scrambling datasets, dataset statistics, 175B samples, a model card, and benchmark-overlap examples.
Official demonstration code implementing the paper's language-model planning and admissible-action grounding workflow.
PIXIU provides FinMA models, FLARE financial evaluation benchmark tasks, and instruction data used as a baseline or source in the paper.
Open-source CAMEL framework for autonomous cooperation among communicative agents using inception prompting and role-play.
Open-source multi-agent collaborative framework associated with MetaGPT, discussed as a representative framework that embeds human workflow processes and SOPs into language-agent collaboration.
Open-source AutoGen framework for creating LLM applications using customizable agents that can be programmed through natural language and code.
Author-maintained repository for tracking LLM-based multi-agent papers and organizing them into streams such as frameworks, orchestration and efficiency, problem solving, world simulation, datasets, and benchmarks.
A GitHub repository made public by the authors to curate papers related to LLM safety.
Repository containing the paper list associated with the survey on LLM-based agents for software engineering.
Repository containing the mathematical problem dataset, evaluation code, and model results associated with the paper.
Repository containing the data and code for the LLM biased reinforcement learning experiments.
The repository contains task implementations for the Iowa Gambling Task, Cambridge Gambling Task, and Wisconsin Card Sort Test, LLM integration utilities, oTree interface components, run scripts, configuration files, and links to analytical code and data.
Repository for TraceLLM, the LLM-based synthetic microservice trace generator and related experimental code.
Official source-code repository for the ICPI algorithm and experiments introduced in the paper.
Repository for the paper's LLM-based anomaly detection implementation, including data-processing scripts, raw data archive, demo scripts for SFT, transfer learning, LoRA, online detection, catastrophic forgetting, and ICL.
A GitHub repository linked by the authors as a resource for recent work on LLMs for ML workflows.
The repository contains the paper's source code in a Jupyter notebook and the RAG knowledge-base text used for the demonstration.
Repository for transformer and foundation models for financial time-series forecasting, including data folders, model implementations, scripts, result files, notebooks, and reproduction instructions.
Fin-LLAMA: efficient fine-tuning of quantized LLMs for finance.
Cornucopia-LLaMA-Fin-Chinese: Chinese finance-oriented LLaMA model referenced by the survey.
An up-to-date resource list for large multimodal agents associated with the survey paper.
The official code repository for experiments and evaluation associated with the Leaky Thoughts paper.
Repository linked by the paper as the available code for the conformal abstention and uncertainty-evaluation work.
Official implementation and experiments for Graph Neural Controlled Differential Equations.
Repository described as the official codebase for the ACL submission titled 'Learning Faster with Better Tokens: Parameter-Efficient Vocabulary Adaptation for Specialized Text Summarization'.
Contains data resources and code for the two-stage training procedure, inference, and evaluation of fine-grained attributed generation.
Repository for reproducing the FAIR paper, including scripts for inferring student mistakes, collecting teacher responses, training distilled student models, and testing accuracy.
Official implementation of the paper's TraceCodegen training framework, including preprocessing, execution, buffer-based self-sampling, training configurations, and inference.
Official implementation and pipeline repository for representation learning in time-domain high-energy astrophysics, including event-file representations, feature extraction, dimensionality reduction, clustering, datasets, encoders, and demonstration notebook.
Repository titled 'BERTOps: Learning Representations on Logs for AIOps' containing data, results, source code, annotated datasets, dataset distributions, and scripts for preparing a pretrained ITOps-domain model.
Code accompanying the paper's limited-expert-prediction learning-to-defer experiments.
Repository containing code for the paper's Robot Language Model approach to grounded task planning.
Contains code for the supervised baseline, trained reward model, PPO fine-tuned policy, model card, and links to the released human-feedback and evaluation data.
Official implementation and pretrained model weights for the CLIP method introduced and evaluated in the paper.
Official experiment code for training LEVER verifiers, executing generated programs, and reproducing the paper's language-to-code evaluations.
CAIL2019 is used as an additional legal case similarity setting involving civil cases such as private lending, IP disputes, and maritime law.
Repository associated with the paper's rationale-augmented dialogue understanding experiments and released resources.
Repository containing the LitSearch dataset and code for constructing or evaluating the scientific literature retrieval benchmark.
Repository named IBM/live-api-bench with code for converting BIRD benchmark SQL queries into API call sequences and producing slot-filling and selection-style benchmark outputs.
Repository for the LiveTradeBench platform/package used to run, monitor, and benchmark LLM-based trading agents across U.S. stock and Polymarket environments.
Contains the LiveVectorLake Python implementation, chunk-level CDC components, Milvus and Delta Lake integrations, query engine, tests, generated test data, benchmark scripts, documentation, and architecture materials.
Official repository linked by the paper for access to the LLaMA model family and associated inference resources.
Repository for the LlamaDuo LLMOps pipeline implementation.
Repository for the LLF-Bench environments, installation instructions, wrappers, and task implementations introduced by the paper.
Repository containing the nudging experimental framework, configuration for nudge types and models, and R analysis code for statistical tests and plots.
Repository associated with the WORKS 2025 Flowcept Agent materials, including software, synthetic workflow, data, analysis code, query set, and prompts referenced for reproducibility.
Flowcept code repository used as the provenance capture and observability foundation for the agent architecture and implementation.
Repository for modernbert_predict_masked.
Repository for medsam_inference.
Repository for flowmap_overfit_scene.
Repository for esm_fold_predict.
Repository for stamp_extract_features and stamp_train_classification_model.
Public repository for TOOLMAKER code and TM-BENCH benchmark.
Repository for pathfinder_verify_biomarker.
Repository for musk_extract_features.
Repository for conch_extract_features.
Repository for uni_extract_features.
Repository for nnunet_train_model.
Repository for medsss_generate.
Repository for tabpfn_predict.
Repository for retfound_feature_vector.
Repository for cytopus_db.
Repository containing the cross-provider validation framework, SEC 10-K test corpus, synthetic financial database, 480 run traces, reproducibility manifests, and setup instructions for release v0.1.0.
Repository containing a Python pipeline for extracting LLM uncertainty features, analysing and selecting features, balancing imbalanced data, training calibrated Ridge and XGBoost meta-models, tuning thresholds, and evaluating cost-aware uncertainty models.
Official implementation of Chameleon, the plug-and-play compositional reasoning system reproduced and structurally analyzed in the paper.
Official repository containing the prompts, outputs, token-count materials, and supporting files for the paper's five experiments.
CAMEL is an open-source multi-agent framework for role-playing and agent collaboration.
Code repository for improving factuality and reasoning through multi-agent debate.
MetaGPT is a multi-agent framework that models a software company using role assignments and SOP-style workflows.
AutoAgents generates different roles for GPTs to form a collaborative entity for complex tasks.
Microsoft AutoGen, a framework for building multi-agent AI applications.
Code repository for Solo Performance Prompting / multi-persona self-collaboration.
AgentVerse provides task-solving and simulation frameworks for multiple LLM-based agents.
ChatDev implements LLM-powered multi-agent collaboration for software development.
Repository connected to AI Scientist-generated papers reported as having passed peer review at an ICLR workshop.
Code repository for MAD, a multi-agent debate framework using large language models.
A curated collection of resources associated with the survey, described by the authors as containing over 200 related papers on agent hallucinations.
Repository containing code, scripts, data folders, requirements, and reproduction instructions for LLM-DSE: Searching Accelerator Parameters with LLM Agents.
Repository reported by the paper as containing the source code and data for the LLM-BLM study.
Repository for reproducing LLM-Explorer experiments and implementation details.
Repository for the paper's LLM-REVal framework, including simulation components for LLM-driven research and review workflows.
Contains the DELEGATE-52 relay runners, direct and agentic model wrappers, prompts, and 52 domain-specific parsers and evaluators.
Official implementation of the Chronos time-series foundation model used for zero-shot inference and sequential fine-tuning.
Configuration files for the CNN-Transformer statistical-arbitrage benchmark replicated in the paper.
Repository directory containing the IPCA, PCA, and Fama-French residual-return datasets used in the paper's backtests.
Official code repository for the paper, containing evaluation scripts, configuration files, a CSV of evaluation results, notebooks, figures, and a modified Lingua training submodule.
Repository for the Needle-In-A-Haystack long-context retrieval test adapted in the paper's fixed- and random-needle experiments.
Official repository for the LLoCO framework and experiments.
Official LongBench repository used for the paper's SingleDoc, MultiDoc, and summarization evaluation; baseline numbers are taken from this repository.
Repository for Local-Splitter, an MCP-compatible and OpenAI-compatible outbound LLM request shim that uses a local small model as a triage layer to reduce cloud token usage.
GitHub repository linked by the paper for LocalEval resources; the repository page inspected stated that available resources were being prepared.
Public repository for LoCoBench-Agent, including the benchmark framework, evaluation setup, data-download instructions, and metrics for long-context software engineering agent evaluation.
Repository containing LogicBench data, evaluation code, and supporting reasoning-chain materials.
Public-facing repository for the paper's KD experiments, configuration files, training code, plotting scripts, and geometric feature analysis utilities.
Repository containing LongReasonArena benchmark data, input-generation utilities, inference code, and evaluation scripts.
Repository containing code, data, and prompts for evaluating LLMs as path planners, including the artifact associated with 'Look Further Ahead: Testing the Limits of GPT-4 in Path Planning'.
Official repository containing the paper's QA and key-value datasets, prompt construction code, data-generation scripts, tests, and experiment instructions.
Repository for LPS-BENCH containing benchmark examples, mock tools, evaluators, prompt templates, schemas, scripts, and the multi-agent case synthesis pipeline.
Repository for the general-purpose MadEvolve code-evolution framework with MAP-Elites, island populations, multiple LLM providers, evaluation backends, and analysis tools.
Repository for the Self-Supervised Audio Spectrogram Transformer used as the architectural and empirical baseline for MAE-AST.
Repository for MaGNet containing Python implementation files for MAGE, hypergraph modules, 2D attention modules, training, datasets/model-weight download links, and backtesting.
Repository containing the authors' simulator and market-making agents for scaled beta policy experiments.
Repository for the FOREC/Cross-Market Product Recommendation baseline used for model parameters and comparisons.
Repository linked by the paper for the efficient cross-market recommendation implementation associated with the proposed market-aware models.
Repository released by the authors for Markov Chain of Thought code and associated research artifacts.
Repository for MASEval, the framework-agnostic multi-agent system evaluation library introduced and evaluated in the paper.
Repository hosting Audio-MAE code and pretrained models for masked spectrogram autoencoding and downstream audio tasks.
Provides code and model assets for MASS pre-training and fine-tuning, including unsupervised and supervised NMT, text summarization, and conversational response generation.
Official repository for the MatPlotAgent framework and MatPlotBench benchmark introduced and evaluated in the paper.
Sibyl System repository included in the selected AutoGen application sample.
AutoTx repository for planning and executing on-chain transactions.
GPT-Academic repository for LLM-assisted academic reading, writing, translation, and code/project analysis workflows.
Composio platform/repository evaluated as a flexible agent application or platform with multiple autonomy-related configurations.
h2oGPT repository for private local GPT-style chat and document interaction.
GraphRag_Ollama repository combining AutoGen, GraphRAG, Ollama, and related tooling.
Langflow platform for building and deploying AI-powered agents and workflows.
Letta platform for stateful agents with advanced memory.
AutoGen open-source framework for building AI agent systems using language models, multi-agent conversations, and tool use.
AutoGen Studio application within the AutoGen repository.
Dream Team repository for building a team of AI agents with AutoGen.
GitHub source for the multiple-choice Truthful-QA variant used in the model-level ranking experiment.
CHALE repository used as a hallucination-evaluation dataset with non-hallucinated, half-hallucinated, and hallucinated answer categories.
Search/RAG infrastructure repository.
Large language model training framework repository.
NLP framework repository.
Glasgow Haskell Compiler repository.
Haskell Cabal build/package repository.
Transformer model library repository.
Property-based testing library repository.
JavaScript DOM implementation repository.
Python spreadsheet/data-analysis tool repository.
Machine-learning library repository.
JavaScript standard library monorepo.
Repository associated with the paper's Time Series Transformer mechanistic interpretability experiments.
Repository for the MedBayes-Lite clinical uncertainty governance layer and associated experimental implementation.
Repository containing the official Python implementation of MEME, including code for argument extraction, mode identification/alignment, and training scripts; the README notes that the dataset used in the project is private due to compliance requirements.
Repository for the MeMemo browser-based HNSW retrieval toolkit, documentation, and RAG Playground example application.
Canonical repository for the MemGPT system, now named Letta, implementing stateful LLM agents with persistent memory and related tooling.
Repository for MemR3, the memory retrieval via reflective reasoning controller.
Official open-source implementation of the MemTools framework introduced and evaluated in the paper.
Repository providing data and code for reproducing analyses, with folders for one-dimensional generated regression, two-armed bandit, real-world regression, and an MMLU benchmark.
Repository containing the MetaOptimize implementation, package files, example usage, and experiment code for CIFAR10, ImageNet, TinyStories, and continual CIFAR100 experiments.
Official MetaTool repository containing ToolE data, tool descriptions and embeddings, scenario lists, prompt templates, model-generation scripts, and evaluation code.
Official repository for the Mind2Web dataset, benchmark processing and evaluation resources, and MindAct fine-tuning and model code.
Repository linked by the paper for the MindWatcher agent framework, models, benchmark resources, and related implementation artifacts.
Repository for MiniCheck code, model usage, synthetic data generation code, benchmark evaluation demo, and links to LLM-AggreFact and Hugging Face model resources.
Repository for the Minions communication protocol enabling small on-device models to collaborate with frontier cloud models.
Repository containing the paper's mKG-RAG implementation and associated resources.
MedicalZooPytorch repository used in the benchmark's OOD training set.
Text classification repository listed as TCL in the benchmark's OOD training repositories.
DeepFloyd IF repository used for multi-modal image tasks.
Deep Graph Library repository used for graph-model tasks.
PyTorch-GAN repository used for image-GAN tasks.
ESM repository used for protein/biomedical tasks.
Public ML-Bench code and benchmark resources released by the paper.
BERT repository used for ML-Bench tasks.
PyTorch Image Models repository used as an OOD evaluation repository.
Grounded Segment Anything repository used as an OOD evaluation repository.
Muzic repository used for music/audio tasks.
OpenCLIP repository used for multi-modal tasks.
vid2vid repository used for video tasks.
OpenDevin agent framework evaluated in ML-Agent-Bench with GPT-4o, GPT-4, and GPT-3.5.
LAVIS repository used for multi-modal tasks.
Stable Diffusion repository used in the benchmark's OOD training set.
Tensor2Tensor repository used in the benchmark's OOD training set.
Time-Series-Library repository used for time-series tasks.
Learning3D repository used for 3D vision tasks.
External-Attention-pytorch repository used for attention-use tasks.
Repository for the benchmark and evaluation resources introduced by the paper.
Official repository containing inference and evaluation code, setup instructions, and links to the released dataset and leaderboard.
Official repository for the MMIR-TCM framework and its associated MedTCM and TDEU artifacts.
Repository for the MMReason benchmark, with citation and evaluation instructions and integration through VLMEvalKit.
Official XiaoMi repository for Mobile-Bench, including Appium/emulator setup materials, API/UI agent code folders, data/result directories, prompts, requirements, and README instructions.
Repository for MobileAgentBench, an automated benchmark for mobile LLM agents with a Python library and default tasks using SimpleMobileTools apps.
Repository named Modalitites-Impact-In-ML under the belgats GitHub account. It was public but empty when accessed.
Public Python repository containing MODE ingestion and inference code, clustering and centroid-routing components, benchmark scripts, evaluation data and logs, tests, and documentation.
Official Model Context Protocol server collection used to identify official and community MCP integrations.
Replication package released by the authors for the MCP server empirical study.
Public repository containing the paper's collected landscape data and implementation examples.
Python/TensorFlow implementation of Robust Log-Optimal Strategy with Reinforcement Learning discussed as a policy-based portfolio optimization approach.
Python/TensorFlow implementation of the PGPortfolio or EIIE-style portfolio reinforcement learning framework discussed in the survey.
GitHub repository for the SYMBA crypto multi-agent reinforcement learning market simulator.
Synthetic train, development, and test files used for the paper's arithmetic argument-extraction experiments.
Auto-GPT is discussed as an autonomous AI application whose main agent, plugins, memory, file access, internet access, code execution, image generation, and oracle-like functions can be represented within the proposed multi-agent graph framework.
BabyAGI is discussed as an AI agent system with task creation, prioritisation, and execution chains that can be modelled as interconnected agents plus a vector-database plugin.
Official source-code repository for the paper, including scalar and two-dimensional consensus experiments, topology configuration, plotting utilities, and experiment execution code.
Official data release containing generated experiment outputs for the agent-count, temperature, and personality analyses.
Repository associated with the experiment in which a fine-tuned Llama 2 7B Chat model attempts to manipulate an LLM overseer or reward model, with RL increasing jailbreak attempts.
Repository associated with the experiment that repeatedly rewrites news articles using LLMs and evaluates degradation in factual accuracy.
Repository associated with the experiment testing whether specialized LLM driving agents fine-tuned on different traffic conventions fail to coordinate when yielding to an emergency vehicle.
Official MH-MoE implementation based on TorchScale and fairseq, including setup and pretraining scripts.
Repository created to support the survey by collecting and categorizing relevant research papers, datasets, application scenarios, framework figures, and future-direction materials on value alignment in agentic AI systems.
Repository containing the Multi-LogiEval data and associated evaluation or reasoning-chain artifacts.
GitHub repository made available by the authors for MMTB.
Repository containing the official implementation of Agentic Predictor for multi-view performance prediction in LLM-based agentic workflows.
Public repository for the MARBLE framework and MultiAgentBench code and data used to develop, test, and evaluate LLM-based multi-agent systems.
Repository identified by the paper as containing the project, data, Python code, ontology model, mapping rules, hyperparameter tuning code for ComplEx, TransE, and DistMult, and code to train and predict using TransE.
A constantly updated paperlist related to the survey topic of multimodality representation learning.
Public repository containing modules that form MRT and utility code used for data piping in the paper's experiments.
Toolkit and repository for creating, sharing, and materializing the natural-language prompt templates used to build P3 and train T0.
Code and instructions for reproducing T0 training, evaluation, inference, and ablation checkpoints.
LASSO is a platform for scalable software code analysis and observation, combining dynamic and static program analysis, code search, N-version assessment, automated test generation, software experimentation, and benchmarking.
Supplemental repository containing accepted submissions, the call for papers, nanopublication files, the questionnaire, and visualization materials for the field study.
Repository containing the nanopublication collection and graph assets used to analyze and visualize the formalization-paper special-issue workflow.
Repository for Nanobench, the template-driven interface used by participants to create formalizations, class definitions, submissions, reviews, responses, updates, and decisions as nanopublications.
Repository for Tapas, the generic triple-store interface used to run template-based SPARQL queries and display submission and review overviews.
Repository providing an open-source NRT-style training recipe built on top of verl, including scripts and configuration for NRT training.
CodRep dataset/competition artifact for code refinement and defect detection.
BigCloneBench dataset for clone detection and related code-understanding tasks.
Code Contests dataset for code generation from competitive-programming problems.
InCoder repository for generative code models.
Description2Code dataset for mapping descriptions to code.
CodeSearchNet dataset covering multiple programming languages for code search and code-language tasks.
APPS benchmark for code generation from programming problems.
Project CodeNet dataset for program understanding, generation, and refinement tasks.
Copilot for Xcode source editor extension integrating GitHub Copilot and ChatGPT-style functions into Xcode.
CodeXGLUE benchmark for multiple code intelligence tasks.
PyCodeGPT/CERT-related repository for Python code generation.
xCodeEval/ExecEval repository for multilingual code evaluation tasks.
HumanEval benchmark for evaluating code-generation models.
WikiSQL dataset for SQL generation/summarization-style tasks in the paper's table.
CONCODE dataset for mapping natural language and program context to code.
Code-LMs repository associated with PolyCoder and code-language-model research.
Official implementation repository for the expanded Natural Language Reinforcement Learning project, including shared NLRL libraries and later Maze, Breakthrough, and Tic-Tac-Toe experiments.
Python code for collecting LLM ranking data, training the scoring model, and training policies with direct-score or potential-difference rewards.
GroundHog is the authors' Theano-based recurrent neural network framework containing the neural machine translation implementation used for the paper.
Repository released by the authors for Multimodal-ZeroShotTM, Multimodal-Contrast, and the multimodal topic modeling experiments.
Repository for NNHedge, containing components for simulated instruments, neural hedging models, data loading, training, and assessment.
Author-linked reference implementation of NSR, including SVD compression, TaskKnowledgeBank logic, allocation policies, dataset builders, experiment runners, and visualization code.
Repository containing code for persona vector generation, the user-study interface, and user-study analysis associated with the neural transparency paper.
Repository containing code, data folders, source files, scripts, requirements, and usage instructions for running CaRing and evaluating ProofWriter, GSM8K, and PrOntoQA experiments.
Repository for the Next-Generation LLM for UAV system, including the main application, short-, medium-, and long-range route-planning utilities, path-planning utilities, control-platform utilities, examples, and integrated data folders.
The repository describes NormCode as a language for auditable multi-step AI workflows, lists core components such as infra, canvas_app, cli_orchestrator.py, documentation, and examples, and links to the arXiv paper.
The fairseq NormFormer example provides architecture flags and causal-language-model training commands corresponding to the paper.
Repository containing code and configuration for training and inference of Qwen3 8B, DeepSeek LLM 7B, and LLaMA3 8B Instruct on financial sentiment datasets, including preprocessing, model configuration, training, inference, and evaluation.
Repository whose README cites the NumHTML AAAI 2022 paper and states that code and data used for the paper are provided.
Repository containing the paper's Off-Policy Corrected Reward Modeling implementation and the code and resulting data for filtering the short Alpaca-Farm setting.
Repository for the ACL 2025 OMGM coarse-to-fine multimodal retrieval and RAG framework.
Repository for the OmniCode benchmark code and data.
Repository associated with the paper's implementation of probabilistic monolithic and model reconciling explanation algorithms.
Repository stated by the paper as containing code to reproduce the experiments for the uncertainty-measure framework.
Repository linked by the paper as the code for the controlled shortest-path reasoning experiments.
Public ReAct repository used as the source code-base for the original ReAct setup and corresponding experiments.
Repository containing the RepoExec benchmark/source code for executable repository-level code generation evaluation and related dependency-utilization tooling.
Repository containing code for the paper, including configurations for BM25 baseline, ReAct agent variants, self-reflection, AutoCodeRover-related settings, and evaluation scripts.
Repository containing the ARC task data and supporting materials for the benchmark introduced and analyzed in the paper.
Repository for the paper's PyTorch implementation, including scripts, prompts, helper functions, and instructions for running memory representation and retrieval experiments.
Official ToolBench repository containing benchmark tasks, action-generator and evaluator code, tests, setup instructions, and evaluation commands.
Repository path for Agent Spec runtime adapters that translate Agent Spec components into framework-specific equivalents for popular agentic frameworks.
WayFlow is presented as the paper's reference runtime for executing Agent Spec components, including native support for Agent Spec Agents and Flows.
Repository announced by the authors for code and data supporting the neutral event graph induction framework.
Open-source code and data repository for the OpenAgentSafety framework and benchmark task suite.
Official implementation repository for OpenAI Gym, the reinforcement-learning environment toolkit introduced and described by the paper.
Repository linked by the authors as the location of all generated code and other answers used in the study.
The open-source implementation of the OpenHands platform introduced and evaluated in the paper.
A modular, parallelized code base for simulating constant-product AMM trading against a CEX, designed for large-scale experiments and extensible market-design variants.
Library of ADMM applications for sparse and low-rank optimization used to test NewADMM.
Huawei Cloud VM-placement traces used in the cloud resource scheduling case study.
Official implementation of OCTree for optimized feature generation for tabular data via LLMs with decision-tree reasoning.
Open-source MA-Gym codebase for evaluating Manager Agents in graph-based multi-agent workflow orchestration.
Open-Finance-Lab AgenticTrading is an open-source experimental playground for LLM-powered trading agents, backtests, paper-trading simulations, reasoning logs, benchmark comparisons, and the FinAgent orchestration subsystem.
Repository for the OrderFusion model and workflow, including package installation, tutorial notebook, data-reading, model optimization, evaluation, and forecast-plotting functions.
Repository for Orla, the library and execution engine for constructing and serving LLM-based agentic workflows with stage mapping, orchestration, and workflow-level memory management.
Official Uber Research repository containing implementations of the original POET and Enhanced POET algorithms.
Public source-code repository for PAMS, the Python-based Platform for Artificial Market Simulations.
Repository for the RedPajama data recipe and corpus from which the paper draws compute-dependent pretraining subsets.
Repository provided by the authors to support reproduction of the Identifier-Organizer-Adapter data synthesis and distillation framework.
Official open-source repository for the PedNStream pedestrian-network simulator, including core LTM modules, scenario data, examples, visualization, tests, and configuration files.
Official ParlAI repository containing the Persona-Chat task infrastructure and associated dialogue-model resources.
Code used to scrape and process the dataset, construct train/test and benchmark datasets, and build and evaluate baseline phishing detection models.
CMBAgent Benchmarks repository used for CAMB tool-grounded precision tasks.
Repository containing the implementation of the Point-M2AE hierarchical point-cloud pre-training framework.
Agent Network Protocol is treated as an open-network agent discovery and collaboration protocol using decentralized identifiers and JSON-LD.
Model Context Protocol is treated as a standardized context-ingestion and tool-invocation protocol relevant to execution-level transitions.
Contains framework code, domain profiles, prompt builders, placeholder QA, deterministic replacement, paired full-resolution benchmark posters, evaluation outputs, runtime and failure audits, prompts, manifests, and configuration records.
Repository associated with the paper's dataset and experimental framework for evaluating LLM-generated scientific reviews against human reviews and post-publication outcomes.
Repository containing code for the paper, including PhraseBank preprocessing, training, and testing scripts for the LLaMA financial sentiment analysis experiments.
Apache Airflow repository; used as an example among ten large open-source industrial projects from which function-summary pairs were sampled.
RxJava repository; used as an example among ten large open-source industrial projects from which function-summary pairs were sampled.
Original StockNet codebase for stock movement prediction from tweets and historical prices.
Repository for the StockNet dataset, containing historical stock-price data and tweet-data structure for stock movement prediction from tweets and historical prices.
Repository reported by the paper as the source code for the LLM-enhanced tweet emotion analysis and stock movement prediction framework.
LLT R package used by the authors to transform the cryptocurrency datasets before classifier evaluation.
A Python benchmark for backtesting prediction-market trading agents using real Kalshi market replay data, included episodes, agent interfaces, simulator configuration, and metrics output methods.
Official code repository for the prefix-tuning method and experiments introduced by the paper.
Official implementation repository for ProAgent, the LLM-based agent system introduced by the paper for Agentic Process Automation.
Public code repository for the trajectory-probing experiments and analyses introduced in the paper.
Repository containing task templates, benchmark-generation scripts, preprocessing, model-prediction configurations, response extraction, and evaluation code for reproducing and extending the ProcBench experiments.
Repository containing code, prompts, and data for reproducing or evaluating Program of Thoughts prompting.
Official repository location for the MBPP programming-problem benchmark introduced by the paper.
Notebook implementing the MathQA-to-Python translation and generation workflow used for MathQA-Python.
Official ProgramBench repository for the benchmark, tooling, usage guide, and baseline evaluation workflow.
Repository containing the progressive multimodal search-agent rollout, image and text tools, retrieval and summarization services, TN-GSPO modifications, training scripts, evaluation scripts, and dataset configuration files.
Repository for the paper's trace-rewriting experiments, including configuration files, source scripts for trace generation, rewriting, distillation and evaluation, prompt-optimization scripts, and pre-generated GSM8K/MATH rewritten trace datasets.
Repository for evaluating LLM agents on real-world coding tasks, with benchmark code, configuration, data folders, unit-test workflow, PyInstruct data link, and PyLlama3 model link.
Repository containing the public implementation and data-processing pipeline associated with the paper.
Repository containing Qlib's code, documentation, and additional platform features beyond those described in the paper.
Python and shell-script repository containing QTMRL or multi-indicator experiment code, non-multi-indicator model scripts, baseline scripts, dependency specifications, and instructions for downloading a related multi-indicator dataset.
Repository containing paper versions, an overview image, README, and Jupyter Notebook implementation for Quantformer.
Repository containing code, data, model implementations, visualisations, and setup instructions for classic and quantile versions of linear regression, BD-LSTM, Conv-LSTM, and ED-LSTM across BTC, ETH, Sunspot, Mackey-Glass, and Lorenz datasets.
Official Quark implementation with separate branches for toxicity unlearning, sentiment steering, and repetition reduction, plus evaluation scripts and released checkpoint links.
Contains the contextual chunker, embedding and reranker training code, Kaggle submission notebooks, official competition data layout, and experiments corresponding to the baseline and final system.
Repository containing the RAGSmith codebase and the datasets used in the study.
Author-maintained repository containing Python and MATLAB implementations and examples for randomizing affine-diffusion models and computing randomized characteristic-function constructions.
Repository associated with Rational Tuning experiments for LLM cascade modeling and threshold optimization.
Alpaca Eval is one of the two central input sets compared in the paper's automatic bencher experiments.
Repository for the paper's released code and data supporting the RealRank automatic LLM ranking analysis.
CHEF is one of the three real-world datasets used to evaluate STEEL.
CrewAI is presented as a Python framework for defining LLM-based agents, tasks, tools, sequential execution, entity memory, and callbacks.
Jupyter notebook implementing a workflow that performs root-cause analysis from a directly-follows graph abstraction and adds evaluation steps such as confidence scoring and reasoning output.
Jupyter notebook implementing the process-mining fairness workflow with protected-group identification and comparison between protected and non-protected cases.
Official Chronos repository for pretrained time-series forecasting models.
Datadog Toto repository for Time-Series-Optimized Transformer for Observability.
Official Google Research TimesFM repository for the Time Series Foundation Model.
IBM Granite TSFM TinyTimeMixer model implementation.
MOMENT repository for a family of open time-series foundation models.
TiRex repository for zero-shot forecasting across long and short horizons.
Moirai/Uni2TS repository for universal time-series forecasting Transformers.
Sundial repository for highly capable time-series foundation models.
Lag-Llama repository for probabilistic time-series foundation forecasting.
Repository for the Re4 Scientific Computing Agent, including source code in Jupyter Notebook and Python script formats and a README describing the rewriting-resolution-review-revision logical chain.
Repository for the ICLR 2023 ReAct prompting paper, including data, prompts, HotpotQA, FEVER, ALFWorld, and WebShop notebooks, plus Wikipedia environment wrappers.
Repository for the paper's training-data synthesis implementation.
GitHub repository associated with the REAL websites, framework, and leaderboard for benchmarking autonomous web agents.
Google Research code for REALM pre-training, document-index refreshing, example generation, released checkpoints, and integration with the ORQA fine-tuning code.
Repository containing the yield-farming simulation environment used for strategy analysis.
Official implementation and release repository for the DramaSR-LRM pipeline, benchmark data structure, training, inference, evaluation, and model checkpoints.
GitHub directory containing REVEAL code, AIGC-text-bank files, configurations, inference scripts, training scripts, and setup instructions.
The ParlAI repository contains the framework, tasks, model access, fine-tuning and evaluation code used to reproduce and extend the paper's chatbot recipes.
Repository released by the authors for ReCode, including the framework and associated research artifacts.
GitHub repository for the benchmark code and evaluation tooling introduced by the paper.
HuggingFace Diffusers repository, used for refactoring diffusion-model implementation files such as UNet and scheduler sources.
Official repository released by the authors with implementations, demos, prompts, and research artifacts for Reflexion experiments.
Repository containing the implementation associated with the regime-aware continual adaptive portfolio-management framework.
Repository associated with the paper's Qwen2.5-0.5B SFT, DPO, RLOO, reward-model, Countdown, and external-verifier experiments.
Code repository for the Relational Representation Distillation implementation reported by the paper.
Public source code for the paper's RPS-based ordinal conformal prediction method and experiments.
The Arena Hard repository is used as the pairwise chatbot evaluation setup; the authors generated additional model outputs and scored them with the compared judges.
The public implementation of RepoAgent, an LLM-powered repository agent for generating, maintaining, updating, and understanding repository-level documentation.
Repository containing RepoGraph code for constructing and retrieving repository graph context, plus integrations with Agentless and SWE-agent and scripts for SWE-bench evaluation.
RepoReviewer repository containing the Python backend, Next.js frontend, screenshots, CLI/API/web workflow, and templates or utilities for future empirical evaluation and annotation.
Official codebase for training ReProbe/UHead-style verifiers, generating annotated reasoning datasets, evaluating ReProbe and baselines, and reproducing benchmark tables.
GitHub repository containing data from the paper's experiments, provided to support reproducibility of the proposed evaluation approach.
Repository containing PopQA training and test data, preprocessing artifacts, generator and reranker LoRA adapters, training and testing scripts, prompt-generation code, and utilities.
Repository for Byzantine Fault Tolerance in LLM-Based Multi-Agent Systems, including pilot experiments, prompt-level confidence probing, hidden-level confidence probing, datasets, pretrained confidence probes, and experiment scripts.
The KILT repository provides the standardized Wikipedia knowledge source used for retrieval in both dialogue benchmarks.
ParlAI is the dialogue research framework in which the paper trains and evaluates all models and through which the RAG-based implementations and pretrained models were released.
The paper's RAG implementation and experiment scripts were ported to and open-sourced within the Hugging Face Transformers repository.
Official companion repository for the survey, containing survey materials and an evolving RAG knowledge base covering papers, datasets, benchmarks, evaluation resources, and toolkits.
Open-Finance-Lab repository for FinRL Contest 2024, linked directly in the paper alongside the contest website.
Official repository released by the authors with code, data, and agent-generated traces for reproducing the study.
Implementation repository used for the Rank1 reasoning-based re-ranker evaluated in the pipeline and query-mismatch experiments.
Repository for the BrowseComp-Plus fixed-corpus deep-research benchmark and its evaluation tooling.
Repository for reproducing or using RFEval, the paper's benchmark for auditing reasoning faithfulness under counterfactual reasoning intervention.
Repository containing the implementation and experiments for the risk-aware GUMDP framework and ERM-MCTS evaluation introduced in the paper.
GitHub repository for the simulator used to quantify uncertainty propagation in AI-augmented systems.
Official Risky-Bench repository containing benchmark datasets, data generation scripts, evaluator components, and evaluation workflows for the paper.
Repository linked by the paper as the code artifact for Variational Alignment with Re-weighting; the repository page was empty when checked.
Repository containing prompts and templates used to instantiate LLM-Modulo for Travel Planner and Natural Plan domains.
A project that crawls WeChat users' profile pictures with usernames and visualizes them; used as the source project for the bug-fixing task.
Repository for serving, training, and evaluating LLM routers, including router types corresponding to the paper such as matrix factorization, similarity-weighted ranking, BERT, causal LLM, and random routing.
Official RouterBench code repository containing data converters, prompt embeddings, predictive and cascading routers, evaluation code, configurations, tests, and visualization utilities.
Repository reported by the paper for the AtomicTranslation code used in the language-to-logic translation experiments.
GitHub repository for a convolutional neural stock-market technical analyser used as the starting point for the paper's proposed model.
Repository for SafeArena, a benchmark for assessing harmful capabilities and safety risks of autonomous web agents.
Repository containing code and prompts used when writing the paper, including agent setup notebooks, safety architecture notebooks, image-generation safety notebooks, requirements, utilities, and unsafe agent test requests.
Google Gemini full-stack LangGraph quickstart prototype used as the basis for a query-decomposition, search, reflection, and answer-finalisation scaffold.
Repository for SafeSearch code, dataset, prompts, assets, and red-teaming configurations.
Official implementation of the Vera safety-testing framework, including taxonomy exploration, case generation, adaptive execution, Vera-Bench, and guard-model fine-tuning components.
DeepSpeed repository containing the distributed training framework and MoE functionality introduced and evaluated by the paper.
A revised Tau2-Bench repository used by the authors because they considered the original benchmark data noisy.
Repository for BIG-bench, from which the paper evaluates 62 tasks covering reasoning, knowledge, social behaviour, and other language-model capabilities.
LLM-Random research codebase containing model-training infrastructure and research configurations associated with the fine-grained MoE scaling-law experiments.
NL2Flow is the automated workflow problem generation and symbolic evaluation framework used to generate planning problems, compile PDDL, and evaluate plans.
NL2FLOW-Runner contains code for running the experiments reported in the paper.
Repository containing canonical tool contracts, interface condition renderers, validation and structured diagnostics, deterministic sandbox executors, run matrix harnesses, structured logging, and aggregate metric scripts.
OpenHands CodeAct is used as one of the agent frameworks evaluated on ScienceAgentBench, alongside direct prompting and self-debug.
Repository containing code and data access instructions for ScienceAgentBench, including benchmark structure, agent/evaluation scripts, and links to benchmark materials.
Repository for the Calo-VQ model used by SciFi to reproduce a calorimeter simulation inference and plotting pipeline.
Repository containing implementation details for the SciFi autonomous agentic scientific workflow framework.
Repository for Reasoning-Reinforced Representation for Search, with release-status information and a linked model artifact.
Public repository containing the SecRepoBench implementation, benchmark metadata, descriptions, harnesses, tools, and scripts for running inference and evaluation.
Provides training code, HH-RLHF data with preference-strength information, and the GPT-4-cleaned validation set.
Official code repository containing the implementation needed to reproduce the GazeReward framework and experiments.
Repository containing code for data generation and risk-control experiments associated with selective conformal classification.
Contains data-processing, SFT, iterative reward-model training and filtering, inference, GPT-4 evaluation, and PPO code for Self-Evolved Reward Learning.
Repository for Self-Evolving GPT; the README states that se_gpt_MAIN.py shows the main workflow of the framework.
Official code and data repository for generating Self-Instruct data, classifying tasks, generating instances, filtering and formatting data, fine-tuning GPT3, and evaluating on the released user-oriented tasks.
Repository containing the original Self-RAG implementation, critic and generator data-creation workflows, retriever setup, training scripts, and short- and long-form evaluation code.
Official repository for the SELF-REFINE framework, containing code, prompts, data, task runners, evaluation scripts, and examples for the paper's studied tasks.
Contains the paper's hybrid-retrieval defense, detection implementations, evaluation and experiment scripts, result files, sanitized examples, corpus setup instructions, figures, and reproducibility documentation.
Contains the pipeline for loading CoQA and TriviaQA, generating answers, clustering semantic similarities, computing likelihoods and uncertainty measures, evaluating AUROC, and reproducing the paper's analyses, together with the released hand-labelled semantic-equivalence data.
Code repository linked by the paper for Sentence-BERT and sentence-transformer models.
The paper's released Sequence Tutor implementation within TensorFlow Magenta, including a checkpointed melody RNN.
Code repository for the ShapG method introduced in the paper.
Self-service snack bar kiosk system repository containing requirements documentation, automated Robot Framework acceptance tests, C4 architecture documentation, ADRs, and traceability audit material.
First iteration for vibe coding in the study; implements a snack kiosk system using React, TypeScript, PostgreSQL, Node, and Express.
Lovable-generated campus-treats prototype used as the unstructured vibe coding example.
Open-source Python framework for building and orchestrating linear deterministic agentic workflows.
Repository containing ready-made example workflows, workflow utilities, and usage material for simpliflow.
Repository containing supplementary materials for the paper, organized into metrics, heatmaps, and equity-line outputs.
Official repository for the skfolio Python library, containing the implementation of the portfolio-optimization and risk-management framework described in arXiv:2507.04176.
Qwen3-Coder repository linked by the paper for the 30B open-weights code-generation model evaluated under NO-SPARK and WITH-SPARK conditions.
Repository for the Smoothie label-free LLM routing method.
Public Python implementation associated with the proposed social recommendation system.
Official Google Research directory containing self-contained prototype notebooks for the Socratic Models applications evaluated in the paper.
Official SPA-Bench repository containing benchmark code, data, framework, model server, pipeline, documentation, and setup files.
Repository associated with the Sparse Logit Sampling paper and Random Sampling KD method; at extraction time the README stated that code would be uploaded after company approvals.
GitHub Spec Kit documentation for spec-driven development, including the staged specification, planning, and tasking workflow used as the conceptual and procedural base for Spec Kit Agents.
Repository containing training and distillation scripts, BigBench Hard and mathematical-benchmark evaluation code, processed-data workflows, prompting notebooks, and notebooks for aligning code-davinci and FlanT5 tokenized outputs.
GitHub repository associated with the SPEECH method and released resources.
A curated repository associated with the paper that organizes efficient architecture papers according to the survey's categories.
Repository for the Semantic Pyramid Indexing FAISS/Qdrant plug-in introduced and evaluated by the paper.
Official KernelBench repository used by the paper to benchmark runtime of PyTorch baselines and LLM-generated kernels.
Repository for the SciBORG manuscript, with setup instructions, framework usage examples, agent construction, and benchmarking guidance tied to the paper release.
Directory containing SciBORG benchmark notebooks and trace notebooks corresponding to the supporting-information experiments.
Repository reported by the paper as the source code for reproducing the pumped-storage DDQN state-representation study.
Repository containing the processed datasets, subject code, static-analysis implementation, machine-learning pipeline, scripts, outputs, feature selection, tuning, cross-validation, and SHAP analyses used in the paper.
Repository containing historical changes in S&P 500 constituents, used to restrict backtest positions to stocks that belonged to the index on each date.
IMPROVER is an operational probabilistic weather forecast post-processing system used to bias-correct, calibrate, threshold, smooth, and blend forecast products.
Contains data-engineering utilities, synthetic-data generation, the three experimental training flows, SFT, DPO and DTFT configurations, and instructions for reproducing the final StatLLaMA path.
Repository for the BESSTIE sentiment and sarcasm classification benchmark for varieties of English.
Repository for the InstruSum instruction-controllable summarization dataset referenced and used as an evaluation target in the paper.
Repository indicated by the paper as containing all code and supplementary materials used for the SV-LSTM hybrid model study.
Repository linked by the authors as the location of the curated dataset used for the stock movement and volatility prediction experiments.
Repository titled MSGCA: Stock Movement Prediction with Multimodal Stable Fusion via Gated Cross-Attention Mechanism, containing code and data folders for the proposed framework.
Repository containing code, configuration files, data-preparation scripts, pretraining routines, and trading scripts for STORM.
Official repository for the StoryScope pipeline, configuration, feature taxonomy, development stories, feature assignments, trained XGBoost models, and reproduction scripts.
Official ECCV 2026 release containing StoryAD-QA annotations, answer keys, evaluation code, and generation and answering prompts; it does not redistribute copyrighted movie media.
Code repository for the LLM-agent Cournot competition simulations and experimental workflow.
The referenced Llama 3 model card is associated with Llama-3-8B-Instruct, one of the local open-source models used in the ranking task.
Code released for the paper, including modular segmentation-attention and syntactic-attention layers and training scripts for translation, question answering, and natural language inference.
Searchat is a local-first semantic search system for AI coding-agent conversations, supporting verbatim, distilled, and cross-layer retrieval over agent transcript histories.
Open-source AI hedge fund project whose structured-summary prompt template, JSON action schema, and next-open execution conventions are adapted for the paper's LLM trading-agent backtests.
Public reference implementation of Context-Aware Decoding, the decoding method adapted by FinCAD for parametric look-ahead-bias mitigation.
Open-source SuperHF training code, reward-model training code, PPO-RLHF baselines, experiments, evaluations, and chart-generation resources.
Repository for the code used in the paper, with directories for the rate, prepare, assess, and pipeline stages.
Companion repository that organizes works on LLM-agent evaluation according to the survey's structure and tracks papers, benchmarks, methodologies, and frameworks.
Framework for evaluating and optimizing agents and models in container environments, discussed as part of emerging standardized cross-environment agent evaluation.
LangChain AgentEvals package for evaluating agent trajectories, including trajectory matching and graph-based evaluation.
HAL harness for centralized and reproducible evaluation across agent benchmarks.
Repository for SWE-agent, the LM-based agent system that attempts to fix GitHub issues using an agent-computer interface and configurable tools.
Repository containing SWE-Bench-CL data, dataset construction scripts, naive and agentic evaluation procedures, and LangGraph/FAISS-based continual-learning agent implementations.
Repository linked by the paper for the SWE-CI benchmark, associated code, and evaluation resources.
Repository for the SWE-EVO benchmark, including benchmark materials and evaluation support for coding agents in long-horizon software evolution scenarios.
Official repository released for the SWE-Lancer benchmark, public Diamond split, code, and evaluation environment.
Repository for SWE-QA-Pro, including evaluation materials for direct and agent modes and links to the paper and benchmark release.
Repository containing experiment code for SynthSAEBench SAE architecture evaluations.
Codebase provided by the authors for reproducing TABCF experiments.
The paper states that this repository contains the reproducibility code for the benchmark of tabular classification methods.
Official TaskBench code and dataset directory within the Microsoft JARVIS repository.
A JSON file encoding TDD principles as governance objects with bibliographic grounding, human-oriented intent, AI-native interpretation, operational constraints, and anti-patterns.
Open-source Python framework extending SMAC with meta-learning and ensemble learning for pipeline automation.
Open-source Python implementation of sequential model-based algorithm configuration for hyperparameter optimization.
Open-source Python Auto-Pipeline tool using genetic programming to optimize tree-structured machine learning pipelines.
Open-source Python framework for automated feature generation from relational datasets.
Open-source Python tool for hyperparameter optimization.
Open-source Keras-based framework for searching deep network architectures using Bayesian optimization and network morphism.
Microsoft open-source toolkit for neural architecture search and hyperparameter tuning across local or cloud execution environments.
Open-source TensorFlow framework for automatically learning neural network architectures and ensembles.
Open LLaMA is the publicly available 13B model used in the paper's instruction-based fine-tuning experiments.
Repository containing TextAtari agent/environment code, translators from Gym-style environments to natural language, prompt/decider modules, manuals, language trajectories, and visualization assets.
Python implementation of the three-stage TextReg pipeline, including RuleBank, gradient purification, semantic edit regularization, the guided optimizer, example execution scripts, and paper-aligned metrics.
Repository accompanying the paper, with runnable TEP pipelines, configurations, benchmark support, and implementation code for deep compound AI system optimization.
Repository for the QA-FEEDBACK, Longformer reward-model, and T5/PPO experimental workflow used to study the reward-model accuracy paradox.
Repository associated with the AI Cosmologist paper, described by the paper as containing code and experimental data and by the repository README as containing configuration files, examples, and best AI-generated code for the Galaxy Zoo and Quijote demonstrations.
Repository linked by the paper for LLM uncertainty decomposition, with folders and scripts related to input uncertainty, decoding uncertainty, model uncertainty, data, models, utilities, and uncertainty scoring.
Repository released by the authors for the agent-to-agent negotiation and transaction benchmark code and data.
Meta Research repository associated with the paper's MAE?WSP pre-pretraining approach.
Public implementation of Progressive Transformers retrained and evaluated to provide a BLEU reference scale.
Implementation repository for the sign-pose VAE variants introduced and evaluated in the paper.
Public implementation of Sign-IDD retrained and evaluated as a non-latent diffusion reference baseline.
Repository containing the paper's transparency materials, including SP-1 AI-usage summary, SP-2 navigation index, SP-3 documentation-adequacy account, SP-4 process documentation, and SP-5 development records.
A framework for training or evaluating agentic systems with reinforcement learning.
A Microsoft framework listed as part of the agentic RL framework ecosystem.
An open-source framework for scaling LLM reinforcement learning.
A repository for verifiable environments used in LLM reinforcement learning.
A benchmark for evaluating agents on real software engineering issues.
A framework for tool-use reinforcement learning with LLM agents.
A web environment benchmark used to evaluate autonomous agents on browser tasks.
Repository for the paper's released data, including generated Code Llama outputs used to support follow-on work on ranking LLM code-generation candidates.
Repository accompanying the paper, containing SPR task materials, generated research outputs, case studies for Agent Laboratory and The AI Scientist v2, and pitfall-detection code.
Open-source implementation of The AI Scientist-v2, one of the two AI scientist systems evaluated in the paper.
Open-source implementation of Agent Laboratory, one of the two AI scientist systems evaluated in the paper.
Repository released by the authors for obtaining and preprocessing decaNLP datasets, training and evaluating models, reproducing experiments, and tracking decaScore progress.
Official repository releasing generated stories and the crowdworker and in-house human-evaluation results used by the paper.
Meta Llama Guard 2 model-card repository path for the safeguard model used to classify whether generated outputs violate predefined safety categories.
Official Google Research codebase for reproducing the paper's prompt-tuning experiments, with training configurations, released prompts, and T5.1.1 LM-adapted checkpoints.
T5 code, evaluation metrics, preprocessing routines, and checkpoint index used by the paper for its base models, data preparation, and reproducibility.
Resources and out-of-domain development data for the MRQA 2019 shared task used in the paper's zero-shot question-answering transfer experiments.
Repository associated with the Test-Driven AI Agent Definition paper, containing code or benchmark artifacts for compiling tool-using agents from behavioral specifications.
GitHub Spec Kit is an open-source toolkit for specification-driven development using commands such as /speckit.constitution, /speckit.specify, /speckit.plan, /speckit.tasks, /speckit.analyze, and /speckit.implement.
Official repository containing code and supporting materials for automated literature collection, deduplication, filtering, review assistance, and the paper's experiments.
Companion code and artifact repository containing analysis scripts, prompt templates, query lists, judge labels, numerical result JSONs, probe configurations, leakage diagnostics, Procrustes tests, and steering analyses.
A GitHub repository collecting papers related to LLM-based agents, linked by the survey as a related-papers resource.
Public repository for the AIDev dataset, schema/CSV files, replication materials, and example notebooks supporting analyses of autonomous coding agents in GitHub PR workflows.
Repository linked by the paper for the reproduction work associated with the study.
Repository containing released evaluation experiments and artifacts for baseline agents and model configurations on TheAgentCompany.
Repository for the TheAgentCompany benchmark environment, tasks, data, and evaluation infrastructure introduced by the paper.
A project associated with the survey's core-competency test framework for LLM evaluation.
Repository containing the DeepFund code used to implement the live fund-investment benchmark, agent workflow, prompts, and evaluation system.
The official Time-MoE repository associated with the paper's architecture and released resources.
Repository containing code, data-processing materials, experiment scripts, results, and plotting notebooks for reproducing the paper's analyses.
Repository containing the tinyBenchmarks Python package, demos, tutorials, and links to tiny datasets for estimating LLM performance from curated small benchmark subsets.
Official repository for the TLOB paper, including code folders for models, preprocessing, data handling, configuration, training scripts, requirements, and backtesting script.
BMTools integrates the paper's evaluated tools and provides an open-source platform for extending foundation models with APIs and for building and sharing tool plugins.
Repository for ToolRoCo, a multi-turn tool-using LLM benchmark for collaborative robotic tasks with Cabinet, PackGrocery, and Sort tasks and four cooperation paradigms.
Official repository containing the ToolAlpaca data, prompts, multi-agent generation code, training scripts, evaluation code, and recorded evaluation outputs.
Hosts the ToolCAD project website and paper-facing artifact page.
Repository currently titled Ziqiao-git/C-World but with README content for ToolGym. It describes ToolGym as an open-world tool-using environment built on 5,571 tools across 204 applications and includes code folders for task creation, tool retrieval, state controller, runtime, and evaluation.
Repository for ToolBench/ToolLLM artifacts, including code, trained models, and demo released by the authors.
Repository for the ToolMisuseBench benchmark implementation, generator, evaluator, and experiment reproduction workflow.
Repository for ToolPRMBench, the benchmark and associated code/data for evaluating process reward models in tool-using agents.
Detectron2 is listed among object detection and image segmentation models/frameworks in the TorchTraceAP application table.
PyTorch Holistic Trace Analysis package referenced as the source of profiling metrics and rule-based trace-event runtime outlier analysis.
Open-source implementation of Self-Supervised Contrastive Pre-Training for Time Series via Time-Frequency Consistency.
GitHub repository for the Light Aircraft Game benchmark, which the paper includes as an application simulator in its benchmark characterization table.
The BosqueLanguage GitHub organisation is identified by the paper as the public location for experimental versions of the AISE-related systems, including the Bosque ecosystem components discussed in the paper.
Repository containing X-FM code and pre-trained models.
Repository containing code for survival-model calibration methods including CiPOT and CSD-style calibration, plus experiment reproduction resources.
A GitHub repository linked by the paper/project page that curates efficient-agent papers and resources corresponding to the survey taxonomy.
Repository containing prompts and queries used in the experiments for the LLM-powered security alert investigation workflow.
A maintained collection of papers, methods, benchmarks, and resources associated with RAG-reasoning and agentic deep-research systems.
Apollo is an open autonomous driving platform used as the representative automated-driving system context for the safety-requirements derivation task.
Code repository implementing the paper's unified MoE compression framework and proposed Expert Trimming methods.
ROS framework for embodied intelligence applications using robot API configuration and LLM calls.
ROS 2 command-line interface extension with LLM support.
Stretch AI system orchestrating skills for language-directed mobile manipulation.
MCP server connecting AI assistants to installed ROS 2 applications and system operations.
Project integrating llama.cpp with ROS 2 to enable LLM inference.
Tool using LLMs to generate ROS codebases from high-level descriptions.
ROS/ROS2 MCP server enabling natural-language commands and monitoring of robot states and sensor data.
ROS-MCP project included among the paper's representative ROS/MCP integrations.
Repository containing the implemented prototype, generated code, and execution traces for the LLM workflow generation experiments.
Repository linked by the authors as the paper list for the survey on reasoning in large language models.
Haystack is the framework used to implement the paper's RAG workload with retrieval and question-answering pipeline behavior.
GitHub repository stated by the paper as the available source code for the multimodal financial forecasting work.
DeepMarket is the official open-source Python framework for LOB market simulation with deep learning. It contains TRADES and CGAN implementations/checkpoints, ABIDES-based simulation components, evaluation utilities, and the TRADES-LOB synthetic dataset.
Repository for TradeTrap, the paper's system-level stress-testing framework for LLM-based autonomous trading agents.
Repository for TradingAgents, the multi-agent LLM financial trading framework introduced and evaluated in the paper.
Repository for BIG-bench, one of the principal benchmark suites used to compare Chinchilla with Gopher across diverse language-model capabilities.
Repository containing released model samples for sampling-based NLP evaluations reported in the paper.
Repository containing reference code for ILF experiments, refinement scoring, reward models, evaluation scripts, and links to the released SLF5K and finetuning datasets; the authors note that excluded data-generation and cleaning steps mean it is not a ready-to-run reproduction package.
Repository for TRAJECT-Bench, including public data, tool definitions, query generation materials, and evaluation scripts for model and ReAct-style agentic tool-use evaluation.
GitHub data source cited for the Stochastic Block Model dynamic graph benchmark used in the experiments.
Official ROLAND repository used by the authors to run the ROLAND baseline five times for MAP/MRR comparison.
Repository containing Python code, prompt files, forex price data, annotated sentiment data, prediction files, and scripts for reproducing prompt runs and comparative results.
Repository released by the authors containing code, generated translations, and human quality assessments for the quality-aware cascaded translation system.
Repository containing Tree of Thoughts code, task implementations, prompts, and logged experimental trajectories.
Official AG2 example repository from which the paper draws four representative MAS applications for case-study evaluation.
Open-source implementation repository for the TrinityGuard framework introduced by the paper.
Repository for TrustAgent code, safety regulations, assets/data, and experiment-running instructions.
A custom library referenced by the paper as the interface through which generated Python code controlled the Boston Dynamics Spot robot.
GitHub repository for the Electricity Transformer Dataset used as ETTh1 and ETTm1 benchmark data in the experiments.
Repository for a research demonstration of the Lean-Agent Protocol, including a frontend, FastAPI orchestrator, Lean worker, policy environment, audit log, and natural-language-to-Lean/back-translation workflow.
Repository containing the authors' implementation of UMoE, the shared-expert architecture that unifies attention-MoE and FFN-MoE modules.
Repository associated with the uncertainty-manipulation attack and confidentiality-preserving audit protocol.
Repository associated with training-dynamics-based selective classification experiments.
Repository associated with experiments and analysis for decomposing the selective-classification gap.
Repository used to reproduce or support the private selective-classification experiments.
The vLLM repository provides the production inference engine and fused MoE execution pipeline into which the authors integrate their activation-sparse routed-expert code path.
Open-source package for replicating experiments, with raw trajectories hosted on Hugging Face, analysis scripts, and data referenced in the paper.
Official SWE-bench experiments repository used to retrieve public agent logs, trajectories, and patch diffs for studied configurations.
Repository containing metadata, analysis data, tools, and code for the cryptocoin correlation analysis.
Official repository containing ITERATER datasets, preprocessing code, intent-classification code, revision-model training and evaluation code, and demonstration materials.
Repository for MAFBench, the unified benchmark suite introduced by the paper for controlled evaluation of multi-agent LLM frameworks.
Repository for ORCA, described as a step toward automating multi-agent system construction from high-level task descriptions using empirical benchmark evidence and cost-aware execution models.
Public Python repository containing data, source code, scripts, tests, results, analysis outputs, and paper figures for the Logic-in-LLMs study.
Repository for the LLM-ification of CHI review, including sampled CHI papers, qualitative codes, metadata, taxonomy images, and supplementary materials.
Open-source implementation of LIME used by the authors to produce local feature-based explanations for the income prediction and biography classification tasks.
Source code for USEagent, the unified software-engineering agent architecture evaluated in the paper.
Source code for USEbench, the unified software-engineering benchmark used to evaluate USEagent and baselines.
Official implementation of CROSS for the paper, including model code, utilities for LLM temporal-chain embeddings, training scripts for temporal link prediction, logs, and dataset-preparation instructions.
Repository collecting materials related to data assessment and selection for language-model instruction tuning.
Repository for the source code used in the paper's statistical-significance benchmark of online regression over multiple datasets.
Repository for the TextFusionHTS framework and experiments reported by the paper.
Official implementation of the precedent-based reaction-plausibility evaluator used as URSA's automated Solv-2 component.
Official implementation of URSA, including route validation, building-block checks, collapsed route variants, ChemCensor scoring, and dataset-level Solv-N metrics.
Repository accompanying USF-MAE with pretraining code, preprocessing notebooks, pretrained checkpoints, figures, and access links for OpenUS-46.
Code repository for D3PO, the paper's reward-model-free direct preference fine-tuning method for diffusion models.
Repository containing code for Using Non-Expert Data to Robustify Imitation Learning via Offline Reinforcement Learning, including simulation pipeline, scripts, installation instructions, and training/evaluation commands.
Repository containing code for cryptocurrency price prediction using RNN-based models and comparison of LSTM, GRU, and Bi-LSTM methods.
Stable Baselines3 is cited as the implementation source for state-of-the-art baseline algorithms such as SAC, TD3, and PPO used in the experiments.
Codebase for the variational-quantum-circuit DDPG/DQN portfolio agents, baselines, training, evaluation, and reproducibility workflow introduced in the paper.
Flask repository used to illustrate semantic, Louvain, and label-propagation clustering over source files.
Python Poetry repository used to demonstrate graph construction, object statistics, issue-driven retrieval, and top-k results.
OWASP IoTGoat is deliberately insecure OpenWrt-based firmware containing vulnerability challenges mapped to the OWASP IoT Top 10.
Official API and supporting resources for running the VirtualHome household simulator and executing activity programs.
Unity source code for constructing VirtualHome environments and translating activity programs into low-level executable character actions.
Repository containing raw collected data, Study 1 to Study 3 folders, preregistration documents, analysis scripts written and executed by the system, analysis outputs, and manuscripts.
Official repository linked by the paper for Visual Semantic Entropy.
Repository containing code associated with W-RAG weak-label generation and retriever fine-tuning experiments.
Repository containing code and data for BadAgents, including poisoned data and code paths for Query-Attack, Observation-Attack, and Thought-Attack experiments.
The GitHub repository hosts the WebArena code, browser environment, evaluation harness, configuration files, Docker environment resources, prompts, scripts, and reproduction materials for the paper.
Repository for the WebShop environment, product and instruction setup, search engine, baseline rule/IL/RL models, tests, and sim-to-real transfer code.
Repository for the ICLR 2025 AI feedback tool that evaluated submitted reviews for vagueness or genericity, possible misunderstanding of the paper, and unprofessional tone, then generated private improvement suggestions.
LLMAgora, the configurable arena used to run two-agent scenarios with public utterances, private reflections, surveys, parameter sweeps, logging, and optional semantic, NLI, and emotion analyses.
Paper-linked GitHub URL for the BigData22 stock-movement dataset; the URL returned 404 during extraction, so repository availability was not verified.
Repository associated with Hybrid Deep Sequential Modeling for Social Text-Driven Stock Prediction and its dataset.
Repository releasing a stock movement prediction dataset from tweets and historical stock prices.
Repository for StockAgent, the LLM-based multi-agent stock trading simulation framework studied in the paper.
A research-grade signal-only decision-support system for cross-sectional ranking of AI-focused U.S. equities with uncertainty quantification, regime-aware deployment gating, PIT-safe data handling, and walk-forward evaluation outputs.
Official implementation of OntoGraphRAG v1.0.0, the experiment harness, scripts for all tables and figures, per-query run logs, GPS replay stores, robustness artefacts, and reproducibility manifests used by the paper.
MM-TSFlib implementation used for Informer, FEDformer, PatchTST, iTransformer, and DLinear backbone experiments.
TimeCMA codebase used as an evaluated aligning-based MMTS comparison.
LeRet codebase used as an evaluated aligning-based MMTS comparison.
Time-LLM codebase used as an evaluated aligning-based comparison model.
Codebase for the Context is Key forecasting benchmark, reproduced to examine LLM performance scaling.
Repository containing an anonymized CAIA evaluator implementation, a benchmark.csv dataset with 178 evaluation questions, evaluation scripts for with-tool and without-tool settings, mock tools, prompts, and dependencies.
Repository containing the compact RLHF pipeline, transition classifier, experiment configurations, analysis scripts, manuscript sources, result tables, figures, examples, and a Gradio-based response-comparison interface.
Repository identified as the code for 'When Routing Collapses: On the Degenerate Convergence of LLM Routers'.
Repository stated by the paper as the location where code and data for AgentDebug will be available.
NVIDIA's transformer inference engine, extended in the paper with DeepSpeed MoE support, TUPE attention, expert routing, quantized MoE computation, and batch pruning.
Repository containing prompts, research ideas, selected outputs, and failure analyses for the four autonomous research attempts, including MARL-idea, SALVO-WM-idea, SDTS-WM-idea, SemEnt-ALGN-idea, and workflow prompts.
Repository for preparing, analyzing, and visualizing survey responses for the 'Will Agents Replace Us?' preprint project, including Python scripts for data preparation, exploratory analysis, inferential analysis, and manuscript figure generation.
Repository for DeepFund, a platform intended to evaluate LLM trading capability across financial markets using a unified environment, multi-agent system, external information ingestion, trading decisions, and arena-style performance presentation.
Repository for Latency Sensitive Benchmarks, including HFTBench and StreetFighter benchmark code and evaluation examples for latency-aware LLM-agent assessment.
Repository linked by the paper for the proposed automated wireless-agent workflow design system.
Repository for WirelessBench, including the wireless tasks and released scoring code.
Framework used for running, managing, and reproducing web-agent experiments on BrowserGym benchmarks.
Open-source Gym-style browser environment for implementing and evaluating web agents with rich observations and action spaces.
Open-source benchmark package for evaluating browser agents on ServiceNow-based knowledge-work tasks.
OpenBMB/WorkflowLLM is the official repository for the WorkflowLLM project, described as a data-centric framework for enhancing LLM workflow orchestration with WorkflowBench and WorkflowLlama resources.
Official implementation artifact for Worldscape-MoE, including training and inference entry points, modality-specific data preparation, validation tools, and optional offline VAE encoding.
Official code repository for training and evaluating Direct Preference Head models, including benchmark evaluation scripts and links to released model checkpoints.
Repository for the paper's XGBoost-based NEPSE log-return forecasting workflow and benchmark outputs.
Public repository containing the cost-aware LLM routing system, training/data-preprocessing components, evaluation and serving pipeline, router tests, and documentation.