Cognaptus DataHub Monitor

GitHub Resources from arXiv Digests

A monitored reference page for GitHub repositories surfaced from arXiv-paper digests, rendered from a machine-generated local data file.

Updated 2026-09-14 20:32:41 +0800 Source: zl_agentr litxr digest extraction GitHub repositories linked from arXiv-paper digests

Tracked Repositories

2045

Unique Papers

1563

Core Fields

paper title, ref_id, GitHub URL

Refresh Mode

local data file written by backend tasks

How To Use This Page

Use it as a lightweight index of implementation assets surfaced from research digests. The page stays defensive: required fields remain visible even when optional metadata is missing.

This page is designed as a refreshable reference surface rather than a hand-maintained article.

The goal is simple: when paper-digest workflows identify linked GitHub repositories, keep them visible in one place with enough context to scan quickly and revisit later.

Repository Index

Search by paper title, arXiv reference, GitHub repository, author, or tags when those fields are available.

Repository containing code, configurations, prompts, and data-processing utilities for according-to prompting and QUIP-Score experiments.

orionw/according-to implementation
Open GitHub

Official repository for the paper containing the in-the-wild prompt corpus, forbidden-question set, ChatGLM evaluator, and semantic-visualization code.

verazuo/jailbreak_llms dataset
Open GitHub

Repository containing code and results for the LLM-based zero-shot tree induction and embedding experiments, with feature-description files referenced for private-dataset feature names.

ml-lab-htw/llm-trees implementation
Open GitHub

Repository containing code and data for replicating the paper's LIME and SP-LIME experiments.

marcotcr/lime-experiments
Open GitHub

Repository linked by the paper as the code for PriorDynaFlow, the proposed a priori dynamic multi-agent workflow construction framework.

L1n111ya/PriorDynaFlow framework
Open GitHub

Repository for the KindsOfReasoning collection and raw outputs/evaluation results for OpenAI instruction-tuned models.

Kinds-of-Intelligence-CFI/KindsOfReasoning dataset
Open GitHub

Repository identified by the paper as containing code for the empirical experiments.

LoryPack/ReferenceInstancesPredictability implementation
Open GitHub

Repository containing the experimental software associated with the proposed causal evaluation framework for deferring systems.

andrepugni/PODS implementation
Open GitHub

Python, notebook, and spreadsheet materials for driver ranking, forced ARMA changes, and related paper experiments.

omhariyadav20/causal_drivers dataset
Open GitHub

EvolveGCN repository data folder referenced as the source for the SBM dynamic graph benchmark.

IBM/EvolveGCN dataset
Open GitHub

Repository containing datasets and code files supporting the paper's DeepSeek-versus-other-LLMs benchmark.

ZhengTracyKe/DeepSeek-and-other-LLMs dataset
Open GitHub

Repository containing accompanying code and resources for the book, including chapter folders and Python-oriented examples for XAI techniques.

Echoslayer/XAI_From_Classical_Models_to_LLMs implementation
Open GitHub

Aegis is an open-source defense system directly evaluated as one of the paper's jailbreak defenses.

automorphic-ai/aegis
Open GitHub

The paper's LLM Jailbreak Attack-Defense Arena implementation, containing attack and defense modules, benchmark data, evaluation utilities, and support for the three tested models.

ltroin/llm_attack_defense_arena benchmark
Open GitHub

LLM Guard is an open-source LLM security toolkit directly evaluated in the paper's defense benchmark.

protectai/llm-guard
Open GitHub

An implementation referenced for reducing GPU all-to-all communication bottlenecks in large-scale MoE execution.

deepseek-ai/DeepEP system
Open GitHub

A GitHub repository associated with the survey that organises resources and representative work on self-evolving AI agents.

EvoAgentX/Awesome-Self-Evolving-Agents
Open GitHub

Repository implementing the LED baselines, evidence-selection scaffold, shorter-context experiments, heuristic evidence baselines, and evaluation workflow reported for QASPER.

allenai/qasper-led-baseline
Open GitHub

Repository containing the EIIE portfolio-management code, configurations, figures, tables, and data used for the cryptocurrency replication and stock-market extension.

jackieli19/PGPortfolio framework
Open GitHub

PGPortfolio (Policy Gradient Portfolio), the authors' source-code repository implementing the paper's deep reinforcement-learning portfolio framework.

ZhengyaoJiang/PGPortfolio
Open GitHub

AIDE-ML: AI-Driven Exploration in the Space of Code, used as the agentic AutoML baseline/interface for observing decisions and artifacts.

WecoAI/aideml framework
Open GitHub

Git repository providing open access to the research datasets supporting the study's findings and replication.

bigrasam/CRNData dataset
Open GitHub

Source code for a minimal MedLog prototype using an OpenAPI-described HTTP REST interface.

mims-harvard/medlog system
Open GitHub

Repository for the financial news dataset used with permission as part of the paper's Bloomberg dataset experiments.

philipperemy/financial-news-dataset dataset
Open GitHub

Open implementation code for the TGN-SEAL framework introduced and evaluated in the paper.

nssajadi/tgn-seal framework
Open GitHub

Public codebase containing the Agent-Driver pipeline, tool library, cognitive memory, reasoning engine, scripts for fine-tuning and inference, prepared data layout, and nuScenes evaluation workflow.

physical-superintelligence-lab/Agent-Driver framework
Open GitHub

The public Lean 4/Mathlib library containing the formalized QAOA and Ising-ring components, the lower-bound and attainability theorems, and the machine-checked proof of residualEnergy_isLeast.

urikol/QuantumOptimization
Open GitHub

Repository released by the authors for the minimal agentic theorem prover and experiment reproduction.

Axiomatic-AI/ax-prover-base framework
Open GitHub

Official code repository for interaction-aware retargeting, retargeted trajectories, Isaac Lab task environments, PPO training, and evaluation scripts for the paper.

yunhaif/regrind
Open GitHub

Repository for the CLINC150 intent-classification and out-of-scope evaluation data used in the experiments.

clinc/oos-eval dataset
Open GitHub

Repository for the StackOverflow short-text intent dataset evaluated by the paper.

jacoxu/StackOverflow dataset
Open GitHub

Repository containing the Banking77 fine-grained banking-intent dataset evaluated by the paper.

PolyAI-LDN/task-specific-datasets dataset
Open GitHub

Repository for a hybrid LSTM, multi-head attention, Gaussian fuzzy-rule, and ARIX local-model forecasting system, including folders for LSTM, Transformer, Neuro-fuzzy, data utilities, model code, and an experiment notebook.

mihaozbot/Fuzzy-transformer system
Open GitHub

R-based analysis and reporting pipeline orchestrated with DVC, with scripts, parameters, environment configuration, intermediate outputs, and instructions for reproducing the reported results.

tmr-crypto/wf_optim_crypto_analysis implementation
Open GitHub

Repository containing code used for the paper, including notebooks and utilities for the daily trading strategy, the CNN model with the new loss, data downloading/processing, baselines, and analysis.

Tony-Guo-1/daily_trading_strategy system
Open GitHub

Repository containing Python, R, and notebook implementations, data files, portfolio outputs, and figures associated with the dynamic stock-recommendation study.

AI4Finance-Foundation/Dynamic-Stock-Recommendation-Machine_Learning-Published-Paper-IEEE implementation
Open GitHub

Repository for Dedupe, the SVM-distance and hierarchical-clustering baseline evaluated against PDDM-AL.

dedupeio/dedupe
Open GitHub

Repository associated with the proposed red-teaming framework, containing generated faithfulness/completeness datasets and prompt materials; the paper states that model configurations, evaluation scripts, standardized judge instructions, attacker prompts, and consensus aggregation code are made available there.

Raed-Mughaus/Red-Teaming-Datasets
Open GitHub

A repository created to keep pace with the fast-moving literature on LLMs and LLM-based agents in science.

ur-whitelab/LLMs-in-science
Open GitHub

Repository released by the author for resources associated with the LLM-agent survey.

xinzhel/LLM-Agent-Survey
Open GitHub

TensorTrade is cited as an open-source package/platform relevant to implementing and adapting RL methods to finance.

tensortrade-org/tensortrade framework
Open GitHub

Official code repository for the paper's Agora implementation and heterogeneous 100-agent demonstration.

agora-protocol/paper-demo
Open GitHub

The GitHub repository contains the self-improving coding-agent implementation introduced by the paper.

MaximeRobeyns/self_improving_coding_agent framework
Open GitHub

FakeNewsNet repository containing fake-news research data resources and crawler tooling for collecting news and related social-media data.

KaiDMML/FakeNewsNet dataset
Open GitHub

Repository containing the resulting mind map, collected notes for each reviewed grey resource, and an online repository of references for further exploration.

SAILResearch/replication-24-harsh-generative-ai-release-readiness-checklist dataset
Open GitHub

The Google Agent2Agent protocol repository, reviewed as a general-purpose inter-agent protocol and used in the paper's comparative use-case analysis.

google/A2A
Open GitHub

Repository for the Web-Agent Protocol, reviewed as a domain-specific inter-agent or human-computer interaction protocol.

OTA-Tech-AI/web-agent-protocol
Open GitHub

Repository for the agents.json specification, reviewed as a domain-specific context-oriented protocol for exposing website capabilities to agents.

wild-card-ai/agents-json
Open GitHub

Companion repository maintained by the authors to track ongoing developments in AI agent protocols.

zoe-yyx/Awesome-AIAgent-Protocol
Open GitHub

Author-maintained companion repository containing the survey's framework materials and a curated, stage-organized list of AI Scientist papers, systems, benchmarks, and applications.

Mr-Tieguigui/Survey-for-AI-Scientist
Open GitHub

An Awesome Data Agents repository linked directly in the paper header, likely used to collect or organize data-agent resources associated with the survey.

HKUSTDial/awesome-data-agents
Open GitHub

JoyAgent is discussed as a proto-L3 system that begins to address predefined-toolset limitations through tool evolution and multi-level thinking.

jd-opensource/joyagent-jdgenie system
Open GitHub

GitHub repository established by the authors as the project page associated with the survey on embodied learning for object-centric robotic manipulation.

RayYoh/OCRM_survey
Open GitHub

Repository linked by the paper as the full list of surveyed papers and summary slides.

junhua/awesome-finance-ai-papers
Open GitHub

Repository for AlphaFin, a benchmark/resource associated with financial question answering and stock prediction.

AlphaFin-proj/AlphaFin dataset
Open GitHub

Repository for R-Judge, a benchmark for safety judgment and risk identification.

Lordog/R-Judge benchmark
Open GitHub

Repository for FinanceBench, a financial question-answering benchmark.

patronus-ai/financebench dataset
Open GitHub

Repository for a Japanese financial language-model benchmark.

pfnet-research/japanese-lm-financial-benchmark benchmark
Open GitHub

Repository for BBT-Fin/CUGE-related Chinese financial language benchmark resources.

ssymmetry/BBT-FinCUGE-Application benchmark
Open GitHub

Repository for FinEval, a Chinese benchmark for financial domain knowledge.

SUFE-AIFLM-Lab/FinEval benchmark
Open GitHub

Repository for PIXIU/FinMA-related financial LLM resources, instruction data, and evaluation benchmarks.

The-FinAI/PIXIU dataset
Open GitHub

Repository for CFBenchmark, a Chinese financial benchmark covering multiple financial NLP tasks.

TongjiFinLab/CFBenchmark benchmark
Open GitHub

Repository for DocMath-Eval, used for evaluating numerical reasoning over text and tables.

yale-nlp/DocMath-Eval dataset
Open GitHub

Official repository for the survey, maintaining a categorized collection of papers, projects, datasets, tools, and related public materials.

OpenDataBox/awesome-data-llm
Open GitHub

A curated repository of papers and benchmarks associated with the survey on reasoning with foundation models.

reasoning-survey/Awesome-Reasoning-Foundation-Models
Open GitHub

Maintained repository containing the survey's paper list, figures, taxonomy-oriented table of contents, and contribution channel for adding omitted work.

UbiquitousLearning/Efficient_Foundation_Model_Survey
Open GitHub

The authors' repository organizes VLM papers, models, datasets, evaluation resources, alignment methods, applications, and challenge areas beyond the static paper.

zli12321/Vision-Language-Models-Overview
Open GitHub

Official repository for the survey, containing the README, poster, and datasets.csv catalog used to share the dataset analysis.

katesanders9/grounded-events
Open GitHub

agentUniverse, a multi-agent ecosystem for autonomous agents.

agentuniverse-ai/agentUniverse framework
Open GitHub

Agno, described as an agentic workflow framework for LLM applications.

agno-agi/agno framework
Open GitHub

Phidata, described as a framework for building multi-modal agents with memory, knowledge, tools, and reasoning.

agno-agi/phidata framework
Open GitHub

Coze, described as an open framework for building agentic applications.

coze-dev/coze framework
Open GitHub

Flowise, described as a drag-and-drop UI to build LLM apps with LangChain.

FlowiseAI/Flowise framework
Open GitHub

LangGraph, described as a stateful multi-actor workflow library for LLM applications.

langchain-ai/langgraph framework
Open GitHub

Dify, described as an open-source LLM application development platform.

langgenius/dify framework
Open GitHub

Microsoft Semantic Kernel, compared as an agent workflow system.

microsoft/semantic-kernel framework
Open GitHub

n8n, described as a fair-code workflow automation platform with UI and integrations.

n8n-io/n8n
Open GitHub

OpenAI Swarm, described as a multi-agent framework by OpenAI.

openai/swarm framework
Open GitHub

Qwen-Agent, a QwenLM agent framework/repository compared in the survey.

QwenLM/Qwen-Agent framework
Open GitHub

Official companion repository that organizes papers relevant to language-model data selection across the training stages and method categories used by the survey.

alon-albalak/data-selection-survey
Open GitHub

Transformer model library used as an inference baseline in the framework comparison.

huggingface/transformers
Open GitHub

Inference and deployment framework used in the paper's W4A16 quantization benchmark.

InternLM/lmdeploy
Open GitHub

LLM serving framework compared for inference and serving throughput, memory management, batching, and scheduling.

ModelTC/lightllm
Open GitHub

NVIDIA inference and serving framework used in the paper's W4A16 quantization benchmark and framework comparison.

NVIDIA/TensorRT-LLM
Open GitHub

An author-linked repository organizing papers on retrieval-augmented generation.

USTCAGI/Awesome-Papers-Retrieval-Augmented-Generation
Open GitHub

GPT Engineer, a software-development agent implementation cited in the engineering application survey and open-source project discussion.

AntonOsika/gpt-engineer implementation
Open GitHub

GPT Researcher, an experimental application that uses LLMs for research-question development, web crawling, source summarization, and aggregation.

assafelovic/gpt-researcher implementation
Open GitHub

AI Legion, an LLM-agent implementation cited in the survey's open-source library and reference set.

eumemic/ai-legion implementation
Open GitHub

LoopGPT, an LLM-agent implementation cited in the survey's open-source library and reference set.

farizrahman4u/loopgpt implementation
Open GitHub

AGiXT, an agent framework implementation cited in the survey as a dynamic AI automation platform.

Josh-XT/AGiXT framework
Open GitHub

DemoGPT, a software-development agent repository cited in the engineering application survey and open-source project discussion.

melih-unsal/DemoGPT implementation
Open GitHub

MiniAGI, an LLM-agent implementation cited in the survey's open-source library and reference set.

muellerberndt/mini-agi implementation
Open GitHub

AgentVerse, a multi-agent collaboration framework referenced among surveyed agent systems and open-source libraries.

OpenBMB/AgentVerse framework
Open GitHub

AgentGPT, an LLM-based autonomous-agent system cited in the survey's open-source library and reference set.

reworkd/AgentGPT implementation
Open GitHub

Auto-GPT, an autonomous LLM-agent implementation included in the construction taxonomy and open-source library discussion.

Significant-Gravitas/Auto-GPT implementation
Open GitHub

SmolModels/developer-style agent repository cited as a software engineering application artifact.

smol-ai/developer implementation
Open GitHub

WorkGPT, a workflow-oriented LLM-agent framework cited as similar to AutoGPT and LangChain.

team-openpm/workgpt implementation
Open GitHub

SuperAGI, an autonomous-agent framework cited in the survey's open-source library and reference set.

TransformerOptimus/SuperAGI implementation
Open GitHub

XLang, an LLM-agent/tool-use framework cited as supporting executable language grounding and interaction with databases, web applications, and physical robots.

xlang-ai/xlang implementation
Open GitHub

Curated repository for tracking LoRA-related research updates and supporting discussion around the survey taxonomy.

ZJU-LLMs/Awesome-LoRAs
Open GitHub

Author-maintained repository established to facilitate ongoing updates and sharing of advances related to the survey's MoE literature coverage.

withinmiaov/A-Survey-on-Mixture-of-Experts-in-LLMs
Open GitHub

Repository for the survey on memory mechanisms of LLM-based agents, including the paper link and visual summaries of the survey sections.

nuster1128/LLM_Agent_Memory_Survey
Open GitHub

A curated reading list accompanying the survey, organized around LLM-agent optimization methods, datasets, benchmarks, and applications.

YoungDubbyDu/Awesome-LLM-Agent-Optimization-Papers
Open GitHub

Repository for the paper's code and experimental workflow comparing LLM self-explanations, human rationales, and post-hoc attribution explanations.

oeberle/self_explanations_human_rationales benchmark
Open GitHub

Repository for CaMeL, a dual-LLM architecture with a safety execution layer and access-control policies for preventing data- and control-flow hijacking.

google-research/camel-prompt-injection
Open GitHub

Repository for JailbreakBench, an evaluation framework for robustness against policy-violating behaviors and adversarial prompts.

JailbreakBench/jailbreakbench benchmark
Open GitHub

Repository associated with IsolateGPT, which isolates application execution and uses structured communication protocols between a planner and dedicated submodels.

llm-platform-security/SecGPT
Open GitHub

Repository for Garak, a security-probing framework that generates and executes adversarial probes to assess LLM vulnerabilities.

NVIDIA/garak
Open GitHub

Repository for NeMo Guardrails, cataloged as a programmable self-reflection/guardrail intervention.

NVIDIA/NeMo-Guardrails
Open GitHub

Repository path listed as the source for the gas station revenue dataset.

bighuang624/DSANet dataset
Open GitHub

Repository listed as the source for daily COVID-19 confirmed and recovered case data.

CSSEGISandData/COVID-19 dataset
Open GitHub

Repository listed as the source for SPMD and VED driving and vehicle energy datasets.

ElmiSay/DeepFEC dataset
Open GitHub

Repository listed as the source for the exchange-rate dataset used in LTSF studies.

laiguokun/multivariate-time-series-data dataset
Open GitHub

Repository listed as the source for daily stock opening price data.

z331565360/State-Frequency-Memory-stock-prediction dataset
Open GitHub

Repository listed as the source for the ETT transformer temperature/load dataset.

zhouhaoyi/ETDataset dataset
Open GitHub

Public repository containing the medical RAG application, evaluation framework, text rechunking pipeline, experimental configurations, and generated evaluation outputs associated with the study.

abdullahmoosa/medrag-research
Open GitHub

OpenAI HumanEval benchmark for evaluating code generation with pass@k metrics.

openai/human-eval dataset
Open GitHub

Self-rewarding reasoning LLM implementation referenced as an example of using model-generated judgments to reduce annotation cost.

RLHFlow/Self-rewarding-reasoning-LLM framework
Open GitHub

AlpacaEval benchmark for automatic evaluation of instruction-following model outputs.

tatsu-lab/alpaca_eval benchmark
Open GitHub

Repository containing the Phase 1 SSL and contrastive encoder code, Phase 2 MoE PPO curriculum, inference and backtest scripts, routing diagnostics, and deployment-related components; the repository states that Phase 3 personalization is proprietary and not released.

rpishehv/PublicFinance-RL framework
Open GitHub

Repository containing code for training and evaluating the Transformer electricity price forecasting model and comparing it against EPF toolbox benchmarks.

osllogon/epf-transformers benchmark
Open GitHub

Repository linked from the paper that mirrors the paper title, abstract, framework figure, table of contents, and compiled relevant works for agent categories and attack types.

OSU-NLP-Group/AgentAttack framework
Open GitHub

Author-provided production-oriented implementation of the A-Mem agentic memory system.

WujiangXu/A-mem-sys
Open GitHub

Author-provided repository for evaluating the A-Mem method and reproducing benchmark experiments.

WujiangXu/AgenticMemory
Open GitHub

Open-source MVP and implementation repository for the AAGATE governance platform introduced by the paper.

kenhuangus/AAGATE
Open GitHub

Repository for ABC-Bench, a benchmark for evaluating whether coding agents can explore repositories, edit code, configure environments, deploy containerized backend services, and pass external HTTP/API integration tests.

OpenMOSS/ABC-Bench dataset
Open GitHub

Optimized Faiss exact-search implementation for Intel CPUs.

architecture-research-group/ae-asplo25-iks-faiss package
Open GitHub

Cycle-approximate IKS simulator parameterized with RTL and memory/interconnect timing.

architecture-research-group/iks_simulator system
Open GitHub

Public repository containing the access-controlled website implementation, agent-side experiment scripts, modified SST/Auth component as a submodule, and experiment logs for delegated access workflows.

asu-kim/agentic-website framework
Open GitHub

Public repository for MCTSr, including scripts for running the supported mathematics benchmarks with LLM inference servers.

trotsky1997/MathBlackBox
Open GitHub

Code and dataset repository for ActivationReasoning, including implementation and materials for replication or extension.

ml-research/ActivationReasoning dataset
Open GitHub

Official repository containing the FLARE code, task configurations, experimental data, retrieval setup, and run instructions.

jzbjyb/FLARE
Open GitHub

Repository containing the ADAGE seven-stage construction pipeline, CAPR-ar/CAPR-am/CAPR-ja benchmark data, model-evaluation code, configurations, and reported evaluation results.

AhmedHajAhmed/adage-benchmark dataset
Open GitHub

Repository linked by the paper for AutoCompressor code and released models.

princeton-nlp/AutoCompressors
Open GitHub

Repository containing data loaders, preprocessing, regime models, custom environments, PPO training pipelines, ablations, evaluation outputs, figures, and notebooks associated with the regime-aware portfolio framework.

GabrielNixon/RegimeAware-PPO framework
Open GitHub

SentiFin is cited as a benchmark dataset for sentiment analysis of Indian financial news headlines and is used to fine-tune the LLaMA 3.2 model.

pyRis/SEntFiN dataset
Open GitHub

An installable Python repository implementing AGoT, AIoT, and GIoT and providing experimental setups, datasets, and result files used to reproduce or inspect the paper's evaluations.

AgnostiqHQ/multi-agent-llm framework
Open GitHub

Repository containing the code developed and used for the study, with a README to support replication and methodology exploration.

stellacydong/rl-cvar-insurance-reserving framework
Open GitHub

GitHub repository identified by the paper as containing implementation details for the Dueling DDQN liquidity-provision method and baseline methods.

HaochenZhang717/Uniswap-v3 benchmark
Open GitHub

Repository stated to include scripts for data preprocessing, model training, performance evaluation, and ablation studies for AMDTL.

mlaurelli/amdtl framework
Open GitHub

Repository containing AMDM implementation code, simulation scripts, evaluation scripts, example data, results files, plots, and the paper materials.

Manishms18/Adaptive-Multi-Dimensional-Monitoring dataset
Open GitHub

Longformer repository for the long-document transformer model evaluated as an additional architecture in the paper.

allenai/longformer implementation
Open GitHub

A continuously updated collection of FFM-related publications, tools, datasets, and resources associated with the survey.

FinFM/Awesome-FinFMs
Open GitHub

The paper identifies this repository as containing representative Python code generated under the tested prompt levels.

VonAugustus/SnakeData implementation
Open GitHub

Author-owned repository for the ICME 2024 ATS paper and its proposed scene-text VQA architecture.

FrankZxShen/ATS
Open GitHub

Repository implementing AFlow's nodes, operators, workflow representation, optimizer, evaluator, benchmark integrations, and scripts for reproducing the paper's workflow-search experiments.

FoundationAgents/AFlow
Open GitHub

Claude-Agent-SDK framework used as one of the scaffold alternatives in the agentic scaffold impact study.

anthropics/claude-agent-sdk-python framework
Open GitHub

Public repository for the AgencyBench benchmark and evaluation toolkit released by the authors.

GAIR-NLP/AgencyBench dataset
Open GitHub

OpenAI-Agents-SDK framework used as one of the scaffold alternatives in the agentic scaffold impact study.

openai/openai-agents-python framework
Open GitHub

Repository for the AGENT KB cross-framework agent memory system introduced and evaluated in the paper.

OPPO-PersonalAI/Agent-KB framework
Open GitHub

Official open-source implementation of Agent Laboratory, including the end-to-end research workflow and source modules for mle-solver and paper-solver.

SamuelSchmidgall/AgentLaboratory
Open GitHub

Official Agent Lumos repository containing code for annotation generation, module training, and benchmark evaluation.

allenai/lumos
Open GitHub

Official WebShop environment used as an unseen interactive task in the Lumos generalization evaluation.

princeton-nlp/WebShop benchmark
Open GitHub

Repository for the Agent Mentor / Agent Analytics open-source observability and analytics platform for agentic AI applications.

AgentToolkit/agent-mentor framework
Open GitHub

GitHub path identified by the paper as the code corresponding to the analytics pipeline used for semantic feature analysis.

AgentToolkit/agent-mentor implementation
Open GitHub

Repository for building task/state knowledge, processing training data, constructing the state knowledge cache, LoRA training of agent and world knowledge models, and evaluating the paper's planning experiments.

zjunlp/WKM
Open GitHub

Repository containing experiment code, response matrices, IRT models, task embeddings, LLM-as-a-judge features, adaptive testing code, and scripts for the new task, new response, new agent, and new benchmark experiments.

dariakryvosheieva/agent-psychometrics implementation
Open GitHub

Repository for the paper's virtual trading arena, ArenaTrader implementation, prompts, code, and data.

wekjsdvnm/Agent-Trading-Arena dataset
Open GitHub

Official code repository implementing AWM pipelines for WebArena and Mind2Web in offline and online settings.

zorazrw/agent-workflow-memory
Open GitHub

Repository for the Agent-as-a-Judge project and DevAI-related evaluation artifacts.

metauto-ai/agent-as-a-judge framework
Open GitHub

GPT-Pilot is one of the three open-source code-generation agentic systems benchmarked in the paper.

Pythagora-io/gpt-pilot framework
Open GitHub

A GitHub repository collecting papers and resources for the survey on Agent-as-a-Judge.

ModalityDance/Awesome-Agent-as-a-Judge
Open GitHub

Repository associated with Agent-R, the paper's iterative self-training framework for training language-model agents to reflect and recover from errors.

bytedance/Agent-R framework
Open GitHub

Repository for AgentBench datasets, environments, and integrated evaluation package.

THUDM/AgentBench dataset
Open GitHub

Official AgentBoard repository containing the benchmark/evaluation framework and associated data resources introduced by the paper.

hkust-nlp/AgentBoard dataset
Open GitHub

Indeed Hiring Lab repository tracking the share of job postings mentioning artificial intelligence.

hiring-lab/AI-Hiring-Tracker dataset
Open GitHub

Google's Agent-to-Agent Protocol repository, referenced as the source for A2A, one of the modern agent communication protocols compared in the paper.

google/A2A implementation
Open GitHub

BlenderMCP repository cited as a GitHub integration for Blender Model Context Protocol.

ahujasid/blender-mcp implementation
Open GitHub

PowerAgent PowerMCP repository cited as a GitHub implementation for power-system simulation software MCPs.

PowerAgent/PowerMCP framework
Open GitHub

Repository hosting the draft proteomics_GROUNDING.md epistemic grounding specification and Appendix A preliminary testing materials.

OmicsContext/proteomics-context framework
Open GitHub

LangGraph is used to illustrate graph traversal with persistent state, checkpoints, controlled cycles, guard nodes, approval steps, and typed state updates for production flow engineering.

langchain-ai/langgraph implementation
Open GitHub

Swarm is used to illustrate specialist-agent handoffs and routines as a controllable coordination pattern complementary to state-machine-style graph orchestration.

openai/swarm implementation
Open GitHub

Open-source repository for the Agent4CT loop, all 26 solver implementations, compact recombination solvers, helical-to-fan rebinning pipeline, and evaluation scripts.

akmaier/Agent4CT benchmark
Open GitHub

Agent-zero repository evaluated as a selected agentic framework in the paper's 22-framework benchmark study.

agent0ai/agent-zero framework
Open GitHub

ANUS repository evaluated as a selected agentic framework in the paper's 22-framework benchmark study.

anus-dev/ANUS framework
Open GitHub

Camel repository evaluated as a selected agentic framework in the paper's 22-framework benchmark study.

camel-ai/camel framework
Open GitHub

CrewAI repository evaluated as a selected agentic framework in the paper's 22-framework benchmark study.

crewAIInc/crewAI framework
Open GitHub

MetaGPT repository evaluated as a selected agentic framework in the paper's 22-framework benchmark study.

FoundationAgents/MetaGPT framework
Open GitHub

Google ADK repository evaluated as a selected agentic framework in the paper's 22-framework benchmark study.

google/adk-python framework
Open GitHub

Source code repository released by the authors for reproducing the benchmark comparison.

GPT-Laboratory/22-Agentic-Framework-Comparison-for-Reasoning-Tasks-across-BBH-GSM8K-and-ARC-Benchmarks implementation
Open GitHub

LangChain repository evaluated as a selected agentic framework in the paper's 22-framework benchmark study.

langchain-ai/langchain framework
Open GitHub

LangGraph repository evaluated as a selected agentic framework in the paper's 22-framework benchmark study.

langchain-ai/langgraph framework
Open GitHub

Mastra repository evaluated as a selected agentic framework in the paper's 22-framework benchmark study.

mastra-ai/mastra framework
Open GitHub

PraisonAI repository evaluated as a selected agentic framework in the paper's 22-framework benchmark study.

MervinPraison/PraisonAI framework
Open GitHub

Autogen repository evaluated as a selected agentic framework in the paper's 22-framework benchmark study.

microsoft/autogen framework
Open GitHub

Semantic-kernel repository evaluated as a selected agentic framework in the paper's 22-framework benchmark study.

microsoft/semantic-kernel framework
Open GitHub

TaskWeaver repository evaluated as a selected agentic framework in the paper's 22-framework benchmark study.

microsoft/TaskWeaver framework
Open GitHub

OpenAI-Agents-Python repository evaluated as a selected agentic framework in the paper's 22-framework benchmark study.

openai/openai-agents-python framework
Open GitHub

Swarm repository evaluated as a selected agentic framework in the paper's 22-framework benchmark study.

openai/swarm framework
Open GitHub

Pydantic-AI repository evaluated as a selected agentic framework in the paper's 22-framework benchmark study.

pydantic/pydantic-ai framework
Open GitHub

Qwen-Agent repository evaluated as a selected agentic framework in the paper's 22-framework benchmark study.

QwenLM/Qwen-Agent framework
Open GitHub

AutoGPT repository evaluated as a selected agentic framework in the paper's 22-framework benchmark study.

Significant-Gravitas/AutoGPT framework
Open GitHub

SuperAGI repository evaluated as a selected agentic framework in the paper's 22-framework benchmark study.

TransformerOptimus/SuperAGI framework
Open GitHub

Upsonic repository evaluated as a selected agentic framework in the paper's 22-framework benchmark study.

Upsonic/Upsonic framework
Open GitHub

Agency-swarm repository evaluated as a selected agentic framework in the paper's 22-framework benchmark study.

VRSEN/agency-swarm framework
Open GitHub

BabyAGI repository evaluated as a selected agentic framework in the paper's 22-framework benchmark study.

yoheinakajima/babyagi framework
Open GitHub

A curated and continuously updated collection of papers and resources organized around the survey's foundational, self-evolving, collective, application, and benchmark taxonomy.

weitianxin/Awesome-Agentic-Reasoning
Open GitHub

Code repository for the Agentic Reasoning framework introduced and evaluated in the paper.

theworldofagents/Agentic-Reasoning framework
Open GitHub

Public repository for Agentic Reward Modeling, including the RewardAgent implementation and materials intended to facilitate further research.

THU-KEG/Agentic-Reward-Modeling system
Open GitHub

A continuously updated collection of relevant studies for the Agentic Web.

SafeRL-Lab/agentic-web
Open GitHub

Official code repository linked by the paper for Agentic-DPO; the current release includes the StableToolBench Qwen3.5-2B training pipeline, canonical step pairs, PPA rendering, hard-negative generation, and trainer code.

Schuture/Agentic-DPO
Open GitHub

Repository for the AgenticPay benchmark and framework, including buyer and seller agents, negotiation environments, examples, metrics, and model backends for LLM-based commerce negotiation.

SafeRL-Lab/AgenticPay dataset
Open GitHub

Repository containing the AgentLAB benchmark code, attack scripts, environments, prompts, data, and usage instructions for evaluating LLM agents against long-horizon attacks.

TanqiuJiang/AgentLAB dataset
Open GitHub

Official open-source repository for the Agentless approach introduced and evaluated in the paper.

OpenAutoCoder/Agentless implementation
Open GitHub

Repository associated with AgentOrchestra; it contains the hierarchical agent runtime and evolving protocol/resource implementation used for general-purpose agent research.

SkyworkAI/DeepResearchAgent
Open GitHub

Repository for the AgentRewardBench library, including tools for running agents, running judges, scoring judgments, loading trajectories, and submitting leaderboard results.

McGill-NLP/agent-reward-bench dataset
Open GitHub

Repository associated with AgentRx, the paper's diagnostic framework and benchmark artifact for AI-agent failure attribution.

microsoft/AgentRx dataset
Open GitHub

Official repository for PASB, including the benchmark data, baseline runners, judge code, audit scripts, and documentation.

henrymao2004/agent-sycophancy dataset
Open GitHub

Hermes-Agent is one of the two stateful personal-agent stacks evaluated by PASB.

NousResearch/hermes-agent
Open GitHub

OpenClaw is one of the two stateful personal-agent stacks evaluated by PASB.

openclaw/openclaw
Open GitHub

A curated GitHub repository associated with the paper that lists and classifies related papers on LLM-based agents in software engineering.

DeepSoftwareAnalytics/Awesome-Agent4SE dataset
Open GitHub

The GitHub repository for AgentScope, the multi-agent platform described in the paper.

modelscope/agentscope framework
Open GitHub

Public implementation of AgentStepper and relevant data for the paper's evaluation.

sola-st/AgentStepper dataset
Open GitHub

GitHub Gist containing the Claude Code implementation prompt for the case summarization by file name microservice.

https:/
Open GitHub

GitHub Gist containing the Case Summarization by Given Case Name Workflow pitch generated by the Planning Agent.

https:/
Open GitHub

Official repository for the AgentVerse framework introduced and evaluated in the paper.

OpenBMB/AgentVerse
Open GitHub

AgentWard prototype repository implementing a lifecycle security architecture for autonomous AI agents with native adaptation to OpenClaw.

FIND-Lab/AgentWard framework
Open GitHub

Agile V skills repository v1.3, described as the composable AI-agent skills library that operationalizes Agile V and includes context-engineering patterns.

Agile-V/agile_v_skills framework
Open GitHub

Repository released by the authors for the robust-transformers implementation used with AGRO.

bhargaviparanjape/robust-transformers framework
Open GitHub

A GitHub CLI extension that exposes a compact interface for reading review threads and submitting inline pull request comments.

agynio/gh-pr-review implementation
Open GitHub

Agyn platform for configuring and orchestrating multi-agent systems with explicit communication, roles, and dedicated sandboxes.

agynio/platform framework
Open GitHub

Repository for the AI Act Evaluation Benchmark, including structured EU AI Act compliance scenarios, QA pairs, documentation, and scripts.

davidath/ai-act-evaluation-benchmark dataset
Open GitHub

Official repository released by the authors with code to reproduce all experimental results, implementations of the proposed simple baselines, and the DSPy joint optimizer.

benediktstroebl/agent-evals
Open GitHub

BabyAGI is used as a representative agentic framework showing how LLMs can be embedded in feedback loops to plan, act, adapt, and manage or prioritize subtasks.

yoheinakajima/babyagi framework
Open GitHub

Repository for the violent-python dataset of security-oriented Python code and natural-language descriptions.

dessertlab/violent-python dataset
Open GitHub

Repository for CodeBERT, the pre-trained programming-language model used as the fine-tuned open-model baseline.

microsoft/CodeBERT implementation
Open GitHub

Repository containing an implementation related to reproducing 'Clustering by fast search and find of density peaks'.

AIReproducibility2018/ClusteringByFastSearchAndFindOfDensityPeaks implementation
Open GitHub

User interface and implementation for customizing the visual framework to empirical studies using AI accuracy, human adherence, and final decision-making accuracy, and for computing reliance components and Q.

jhnnsjkbk/accuracy-reliance implementation
Open GitHub

Official open-source repository for Denario, a modular multi-agent system for scientific research assistance that supports idea generation, literature checking, planning, code execution, analysis, paper drafting, and review.

AstroPilot-AI/Denario
Open GitHub

A GitHub Action that reviews pull requests by sending changed file contents to ChatGPT/OpenAI and posting review comments back to the PR.

agogear/chatgpt-pr-review system
Open GitHub

Repository containing the paper's algorithm code and a link to the associated QuantConnect backtest artifact.

tiagomonteiro0715/AI-Powered-Energy-Algorithmic-Trading-Integrating-Hidden-Markov-Models-with-Neural-Networks implementation
Open GitHub

Repository released by the authors for the Writing Quality benchmark, WQRM models, data, and associated evaluation or editing code.

salesforce/creativity_eval
Open GitHub

GitHub repository linked by the paper for example Jupyter notebooks and materials related to AIDev and AI teammates in SE 3.0.

SAILResearch/AI_Teammates_in_SE3 dataset
Open GitHub

Repository for the AIRepr / LLM data-science reproducibility experiments and code released by the authors.

qunhualilab/LLM-DS-Reproducibility framework
Open GitHub

Repository stated by the paper as the public code and data release for AirQA.

OpenDFM/AirQA dataset
Open GitHub

Repository for the AISysRev web application, an LLM-based tool for title-abstract screening that imports CSV metadata, applies inclusion/exclusion criteria using multiple LLMs, supports manual screening with LLM guidance, and exports results to CSV.

EvoTestOps/AISysRev system
Open GitHub

Repository containing ALFRED dataset tooling, data-generation and replay code, Seq2Seq baseline training/evaluation code, documentation, and benchmark submission utilities.

askforalfred/alfred dataset
Open GitHub

Official ALFWorld repository containing the aligned environments, downloadable game/PDDL assets, pretrained components, agents, and training/evaluation code.

alfworld/alfworld implementation
Open GitHub

A curated repository accompanying the paper that organizes embodied-AI papers and resources covered by the survey.

HCPLab-SYSU/Embodied_AI_Paper_List
Open GitHub

Contains the source code, dataset-construction scripts, dependency documentation, and result-generation pipelines associated with the paper.

brains-group/FLARKO framework
Open GitHub

Repository containing code and configurations for training and evaluating AlphaDPO.

junkangwu/alpha-DPO implementation
Open GitHub

Official Alympics repository containing the framework code, Water Allocation Challenge experiment code, requirements, prompts/resources, and related implementation materials.

microsoft/Alympics
Open GitHub

Repository containing the AMVICC prompt CSV, model-evaluation scripts, image-generation scripts, result examples, and utilities used to reproduce or extend the benchmark experiments.

AahanaB24/AMVICC benchmark
Open GitHub

Repository containing a practical implementation of the online AdaVol recursive volatility-prediction method.

nicklaswerge/AdaVol package
Open GitHub

Public repository containing experimental code supporting the TDQN trading results reported in the paper.

ThibautTheate/An-Application-of-Deep-Reinforcement-Learning-to-Algorithmic-Trading benchmark
Open GitHub

Official repository for the paper's uncertainty analysis, ConfiLM resources, and Olympic 2024 data release.

hasakiXie123/LLM-Evaluator-Uncertainty
Open GitHub

Semantic Kernel is described as a modular, plugin-based framework for integrating LLM and agent capabilities into enterprise and software systems.

microsoft/semantic-kernel framework
Open GitHub

Swarm is described as a lightweight multi-agent interface framework for experimenting with multi-agent coordination.

openai/swarm framework
Open GitHub

BabyAGI is described as an experimental framework for autonomous task planning and iterative execution through a self-improving task loop.

yoheinakajima/babyagi framework
Open GitHub

Author-provided Deformable DETR code repository used for the reconstruction-based SSL evaluation.

gokulkarthik/Deformable-DETR implementation
Open GitHub

Author-provided DETR code repository used for the proposed SSL task experiments.

gokulkarthik/detr implementation
Open GitHub

Official SWE-bench experiment repository used by the authors as the source of execution logs, generated patches, and pass/fail statuses for the studied tools.

SWE-bench/experiments dataset
Open GitHub

Code repository for ACEFormer, the attention-based stock forecasting system introduced by the paper.

DurandalLee/ACEFormer implementation
Open GitHub

Repository for the Online-Mind2Web benchmark and associated evaluation artifacts introduced by the paper.

OSU-NLP-Group/Online-Mind2Web dataset
Open GitHub

Google Research repository containing Vision Transformer code, fine-tuning support, and released pretrained models associated with the paper.

google-research/vision_transformer implementation
Open GitHub

GPT Engineer is described as an AI coding agent that can generate an entire codebase from a prompt and ask clarifying questions.

AntonOsika/gpt-engineer implementation
Open GitHub

MetaGPT is described as a multi-agent framework assigning different GPT roles to form a collaborative software entity for complex tasks.

geekan/MetaGPT framework
Open GitHub

AgentGPT is described as a framework for rapidly configuring and deploying autonomous AI agents.

reworkd/AgentGPT framework
Open GitHub

Multi-GPT is described as an experimental multi-agent system in which expertGPTs collaborate, communicate, and use short- and long-term memory.

sidhq/Multi-GPT system
Open GitHub

Auto-GPT is described as an early example of GPT-4 running fully autonomously by chaining LLM thoughts to achieve user-set goals.

Significant-Gravitas/Auto-GPT implementation
Open GitHub

SuperAGI is described as a developer-centric open-source framework for building, managing, and running autonomous AI agents.

TransformerOptimus/SuperAGI framework
Open GitHub

BabyAGI is described as a task-driven autonomous AI agent that builds and prioritises tasks toward an overall goal.

yoheinakajima/babyagi implementation
Open GitHub

A curated repository of reentrancy attacks from which the paper draws 13 sufficiently isolated real-world exploit contracts for evaluation.

pcaversaccio/reentrancy-attacks dataset
Open GitHub

Official implementation of the paper's lightweight contamination detector, including search scripts, contamination reports for six benchmarks, model predictions, clean-dirty comparison code, and visualizations.

liyucheng09/Contamination_Detector
Open GitHub

Repository containing code to generate school-level mathematical question-answer pairs and access the released Mathematics Dataset used for the paper's benchmark experiments.

deepmind/mathematics_dataset dataset
Open GitHub

The report states that EnglandCovid is a PyTorch Geometric Temporal dataset and cites this GitHub repository as reference [9].

benedekrozemberczki/pytorch_geometric_temporal dataset
Open GitHub

Repository referenced by the paper for more details on the EnglandCovid dataset conversion and temporal neural network analysis.

verma-rishu/Analysis_TNNs implementation
Open GitHub

A repository of DAN-style jailbreak prompts represented in the paper's attack-source comparisons.

0xk1h4n/DAN-Jailbreak
Open GitHub

A system-prompt leak collection used to inform the paper's adversarial prompt pool.

asgeirtj/system_prompts_leaks
Open GitHub

A public collection of GPT super-prompts used as part of the paper's heterogeneous attack-source pool.

CyberAlbSecOP/Awesome_GPT_Super_Prompting
Open GitHub

An implementation reference for the self-examination or self-defence mechanism evaluated by the paper.

poloclub/llm-self-defence
Open GitHub

A prompt-injection defence framework used as the reference point for vector-based detection and filtering.

protectai/rebuff
Open GitHub

A curated jailbreak-prompt repository used to broaden the study's community-sourced attack collection.

Techiral/awesome-llm-jailbreaks
Open GitHub

A defence repository referenced for the voting and structured validation approaches discussed and evaluated.

theshi-1128/llm-defence
Open GitHub

A prompt-hacking resource used as one of the repositories from which attack examples were collected.

TrustAI-laboratory/Learn-Prompt-Hacking
Open GitHub

A prompt-hacking collection used as an external source of attack patterns.

yunwei37/prompt-hacker-collections
Open GitHub

Repository released with the paper containing the additional human reference translations and related analysis resources.

facebookresearch/analyzing-uncertainty-nmt
Open GitHub

The fairseq-py toolkit supplies the pretrained convolutional sequence-to-sequence models used as the main objects of analysis.

facebookresearch/fairseq-py
Open GitHub

A curated repository for the paper's agentic-memory survey, organised around the taxonomy introduced in the manuscript and intended to be updated as the field evolves.

FredJiang0324/Anatomy-of-Agentic-Memory
Open GitHub

Official repository for the paper, containing a CLAIR preference-generation notebook, cached results, documentation, and links to associated data.

ContextualAI/CLAIR_and_APO
Open GitHub

Repository for the AndroidWorld environment, task suite, evaluation logic, and associated experiments introduced by the paper.

google-research/android_world dataset
Open GitHub

Official API-Bank code and data directory containing implemented APIs, datasets, database initialization resources, evaluators, simulators, examples, and demo code.

AlibabaResearch/DAMO-ConvAI dataset
Open GitHub

GitHub repository for the DLS-DDPG implementation used or released with the paper.

Hisato-Komatsu/DLS-DDPG implementation
Open GitHub

Repository for AppWorld Engine and AppWorld Benchmark, including the simulated app environment, benchmark tasks, and evaluation infrastructure.

stonybrooknlp/appworld dataset
Open GitHub

Repository for APTBench, the benchmark for evaluating agentic potential of base LLMs during pre-training.

TencentYoutuResearch/APTBench dataset
Open GitHub

Repository identified by the paper as the code for the Arabic tool-calling benchmark work.

kubrak94/gorilla dataset
Open GitHub

AutoResearchClaw is cited as an example of the scenario-verticalized or research-oriented pattern and as the only pipeline/stage subagent case.

aiming-lab/AutoResearchClaw framework
Open GitHub

deer-flow is cited as an example of the scenario-verticalized or research-oriented pattern.

bytedance/deer-flow framework
Open GitHub

cline is included as a corpus project and cited as an example of tool-based delegation.

cline/cline system
Open GitHub

docker-agent is used as a representative orchestration-oriented project combining orchestrator-worker structure, hybrid context, MCP-first tooling, and advanced execution controls.

docker/docker-agent system
Open GitHub

fast-agent is used as an illustrative project showing tool delegation, hybrid context handling, and MCP-first tool registration.

evalstate/fast-agent framework
Open GitHub

Gemini CLI is included as one of the official or first-party products from major AI companies in the corpus.

google-gemini/gemini-cli system
Open GitHub

deepagents is cited as an example of the scenario-verticalized or research-oriented pattern.

langchain-ai/deepagents framework
Open GitHub

Mistral Vibe is included as one of the official or first-party products from major AI companies in the corpus.

mistralai/mistral-vibe system
Open GitHub

Kimi CLI is included as one of the official or first-party products from major AI companies in the corpus and is named as a balanced CLI representative.

MoonshotAI/kimi-cli system
Open GitHub

nullclaw is cited as an example of the Enterprise Full-Featured pattern.

nullclaw/nullclaw framework
Open GitHub

codex is included as an official product/public-evidence case in the corpus and appears in the complete project list.

openai/codex system
Open GitHub

OpenClaw is analyzed as a corpus project and cited as an example of multi-level recursive or balanced CLI-style infrastructure depending on context.

openclaw/openclaw framework
Open GitHub

OpenHands is included in the 70-project corpus and used as an example of event-driven, hybrid-context, enterprise-tool, advanced-safety infrastructure.

OpenHands/OpenHands framework
Open GitHub

agentpool is used to illustrate event-driven subagent architecture with hierarchical context, registry tooling, and enterprise safety classification.

phil65/agentpool framework
Open GitHub

Qwen Code is included as one of the official or first-party products from major AI companies in the corpus.

QwenLM/qwen-code system
Open GitHub

openfang is cited as an example of the Enterprise Full-Featured pattern.

RightNow-AI/openfang framework
Open GitHub

Repository containing three customer-service chatbot variants generated from different single prompts, with methodology, prompts, and source code for the paper's vibe-architecting demonstration.

phomarkon/vibe-architecting-case-study implementation
Open GitHub

GitHub repository for the ARDNS-FN-Quantum code, including the core script and interactive analysis or visualization notebook described by the paper.

umbertogs/ardns-fn-quantum framework
Open GitHub

Repository containing NeuLR and the paper's appendix or supporting resources for benchmark construction and evaluation.

DeepReasoning/NeuLR dataset
Open GitHub

LongBench is used as the main benchmark source for most datasets and for baseline evaluation settings.

THUDM/LongBench dataset
Open GitHub

GitHub repository for the Cross-Attention-only Time Series transformer introduced in the paper.

dongbeank/CATS benchmark
Open GitHub

Repository containing the code implementation for evaluating language models using the paper's framework and generating visualizations; described as configurable for other HuggingFace models, metrics, task types, domains, and reasoning types.

neelabhsinha/lm-application-eval-kit benchmark
Open GitHub

Official repository for MMStar containing the evaluation code and dataset associated with the paper.

MMStar-Benchmark/MMStar dataset
Open GitHub

Repository containing code for ArgEval, the paper's argumentation-based LLM decision-support framework.

adamdejl/argeval framework
Open GitHub

The core open-source implementation repository for the Ark robotics framework introduced by the paper.

Robotics-Ark/ark_framework
Open GitHub

Salesforce repository containing the Art_or_Artifice project folder and associated creativity-evaluation resources.

salesforce/creativity_eval benchmark
Open GitHub

Code repository linked by the paper for the ART framework and its language-program implementation.

bhargaviparanjape/language-programmes implementation
Open GitHub

Companion repository assembled by the authors to collect the papers surveyed in the review and link associated code when available.

dengjianyuan/Survey_AI_Drug_Discovery
Open GitHub

Repository containing Python implementation files and supplementary materials for detecting hallucinations with specialized model divergence.

ACMCMC/ask-a-local implementation
Open GitHub

Repository reported by the authors as containing the code and data for Ask an Expert / BBMHReasoning experiments.

QZx7/BBMHReasoning implementation
Open GitHub

Repository containing code and scripts for the ASPIRE paper, including captioning, GPT-4 prompt use, spurious object detection, diffusion fine-tuning, image generation, and classifier training workflows.

Sreyan88/ASPIRE framework
Open GitHub

Repository released to support reproduction and continued evaluation of prompt-injection attacks and defenses studied in the paper.

sherdencooper/prompt-injection
Open GitHub

Repository implementing Wanda/SNIP-based pruning, safety-utility set difference, ActSVD low-rank removal, orthogonal-projection rank disentanglement, zero-shot utility evaluation, and attack evaluation for Llama2-chat models.

boyiwei/alignment-attribution-code
Open GitHub

Repository containing the ARMT implementation and scripts for training and evaluating the architecture studied in the paper.

RodkinIvan/associative-recurrent-memory-transformer implementation
Open GitHub

Official repository for ATACompressor, containing pretraining, finetuning, inference, preprocessing, configuration, and LoRA implementation code.

Cocobalt/ATACompressor
Open GitHub

Tensor2Tensor repository containing the code used by the authors to train and evaluate the Transformer models.

tensorflow/tensor2tensor
Open GitHub

Repository linked by the authors as the code for the attention-retention continual-learning framework.

zugexiaodui/AttentionRetentionCL framework
Open GitHub

Repository released by the authors containing system responses and their human and automatic attribution ratings for the Attributed QA study.

google-research-datasets/Attributed-QA
Open GitHub

Official Audio-Omni repository containing the model implementation, inference code, API, prompts, interface, configuration, and setup instructions.

ZeyueT/Audio-Omni
Open GitHub

Repository linked from the arXiv v1 project page and identified in its README as the official repository for the AudioDER paper, with dataset access and construction resources.

anonycode26/AudioDER dataset
Open GitHub

Author-referenced sample dataset of synthetic identification-document images covering five document types.

meetsandesh/identification_document_dataset dataset
Open GitHub

Author-referenced repository for generating synthetic document images used in the document identification and information extraction experiment.

meetsandesh/synthetic_document_generator dataset
Open GitHub

Repository containing the implementation of Document Augmentation for dense Retrieval (DAR).

starsuzi/DAR implementation
Open GitHub

Repository identified by the paper as the official source-code location for the Auto-ADMET method.

alexgcsa/auto-admet system
Open GitHub

Repository containing the Auto-FP benchmark resources, including code, datasets, meta-features, and comprehensive experimental results associated with the paper.

AutoFP/Auto-FP dataset
Open GitHub

Official AutoAct repository containing code and prompts for self-instruct data generation, automatic tool selection, trajectory synthesis and filtering, LoRA-based self-differentiation, and group-planning evaluation.

zjunlp/AutoAct
Open GitHub

Source-code repository for the AutoFlow framework introduced by the paper.

agiresearch/AutoFlow framework
Open GitHub

Repository for the Autoformer model and experiments introduced in the paper.

thuml/Autoformer benchmark
Open GitHub

The paper's AutoGen framework repository for building LLM applications via multi-agent conversations.

microsoft/autogen framework
Open GitHub

Repository for Bias Identification Test in Sentiments, including data files and templates for probing sentiment and toxicity models for bias; the paper specifically uses the disability facet.

PranavNV/BITS dataset
Open GitHub

Repository for the ADAS codebase introduced by the paper, including the Meta Agent Search implementation and experimental framework.

ShengranHu/ADAS framework
Open GitHub

Repository for ECLAIR, described as enterprise-scale AI for workflows, including code and scripts for the paper experiments and hospital workflow demo materials.

HazyResearch/eclair-agents framework
Open GitHub

Repository containing folders for datasets, models, plots, pretrained models, additional tests, requirements, and a README with benchmark results for AutoML-DC.

MilanShao/AutoML-for-Multi-Class-Anomaly-Compensation-of-Sensor-Drift dataset
Open GitHub

Open-source implementation of the paper's three-task validation case study, including Python orchestration, generated Prolog programs, Prolog templates, sample-data creation, and logic execution.

cecilpang/autobus-paper
Open GitHub

Official repository for the AG3D museum model, exploration data, and PointNav++ benchmark artifacts.

aimagelab/ag3d
Open GitHub

Official implementation and pretrained models for the impact-driven exploration method developed in Chapter 3.

aimagelab/focus-on-impact
Open GitHub

Official LoCoNav implementation and models for deploying a Habitat-trained navigation system on a real LoCoBot.

aimagelab/LoCoNav
Open GitHub

Official artifact for the Spot the Difference task, manipulated-map dataset, training code, and pretrained models.

aimagelab/spot-the-difference dataset
Open GitHub

Supplementary repository containing datasets, generated manuscripts, run records, coding runs, and supplemental appendix materials used to support the paper's evaluation.

rkishony/data-to-paper-supplementary dataset
Open GitHub

Code implementation of the data-to-paper framework for backward-traceable AI-driven scientific research.

Technion-Kishony-lab/data-to-paper framework
Open GitHub

Hosts the working paper, data.json, figures, tables, and an interactive visualization of Top-N portfolio performance and alpha concentration.

mapledust0/AI-Stock-Nowcasting dataset
Open GitHub

Idea2Paper repository for connecting research ideation to paper-oriented production.

AgentAlphaAGI/Idea2Paper
Open GitHub

Aider repository for agent-assisted code editing and executable software workflows.

Aider-AI/aider
Open GitHub

OpenScholar repository for literature-grounded scientific question answering and knowledge synthesis.

AkariAsai/OpenScholar
Open GitHub

Tongyi DeepResearch repository for agentic deep-research workflows.

Alibaba-NLP/DeepResearch
Open GitHub

OpenHands repository for general software-agent execution over files, tools, and runtimes.

All-Hands-AI/OpenHands
Open GitHub

GPT Researcher repository for automated deep-research report generation.

assafelovic/gpt-researcher
Open GitHub

ScienceClaw repository for scientific-agent workflow orchestration.

beita6969/ScienceClaw
Open GitHub

DeerFlow repository for deep-research workflow orchestration.

bytedance/deer-flow
Open GitHub

PaperQA2 repository for paper-grounded question answering and scientific synthesis.

Future-House/paper-qa
Open GitHub

AI-Researcher repository for an integrated literature, experimentation, and reporting workflow.

HKUDS/AI-Researcher
Open GitHub

autoresearch repository for a steerable software-native AI research loop.

karpathy/autoresearch
Open GitHub

Open Deep Research repository for recursive literature research and report construction.

langchain-ai/open_deep_research
Open GitHub

Agent Laboratory repository for a human-piloted multi-agent computational research workflow.

SamuelSchmidgall/AgentLaboratory
Open GitHub

STORM repository for retrieval-grounded, multi-perspective long-form synthesis.

stanford-oval/storm
Open GitHub

SWE-agent repository for repository-level software-agent execution.

SWE-agent/SWE-agent
Open GitHub

ARIS repository for long-running AI research orchestration and execution.

wanshuiyin/Auto-claude-code-research-in-sleep
Open GitHub

ResearchClaw repository for reusable research-agent workflow infrastructure.

ymx10086/ResearchClaw
Open GitHub

Official repository for AvaTaR with implementation, configurations, data preparation, retrieval-task scripts, and baseline/evaluation code.

zou-group/avatar dataset
Open GitHub

Official source repository for the AVISE framework and the Red Queen SET introduced and evaluated in the paper.

ouspg/AVISE
Open GitHub

Official AxQM repository containing the benchmark task statements, finite-dimensional quantum-mechanics library, Mathlib fork, per-task proof-length estimates, task-dependency ledger, and grading and submission scripts.

Axiomatic-AI/AxQM benchmark
Open GitHub

Repository for the summarize-from-feedback / TL;DR data used as one of the paper's two primary preference-learning evaluation settings.

openai/summarize-from-feedback dataset
Open GitHub

Repository for the BacktestBench benchmark and AutoBacktest implementation, including project folders for AutoBacktest, datasets, figures, tables, environment setup, database setup, and reproduction scripts.

jensenw1/BacktestBench dataset
Open GitHub

Repository for the BadAgent attack on LLM agents.

DPamK/BadAgent implementation
Open GitHub

Repository for the BagStacking implementation provided by the authors within a Scikit-learn API framework.

SeffiCohen/BagStacking package
Open GitHub

AgentGPT is analysed as a general-purpose autonomous LLM-powered multi-agent system with user-guided alignment in selected aspects such as decomposition, agent generation, and resource utilization.

reworkd/AgentGPT framework
Open GitHub

Auto-GPT is analysed as a general-purpose autonomous LLM-powered multi-agent system with autonomous goal decomposition, task action management, and resource utilization.

Significant-Gravitas/Auto-GPT framework
Open GitHub

SuperAGI is analysed as a general-purpose autonomous LLM-powered multi-agent system with some user-guided alignment options for agent-related and resource-related aspects.

TransformerOptimus/SuperAGI framework
Open GitHub

BabyAGI is analysed as a general-purpose autonomous LLM-powered multi-agent system with a profile similar to Auto-GPT across many assessed aspects.

yoheinakajima/babyagi framework
Open GitHub

The fairseq BART directory provides the implementation interface, released BART checkpoints, and task-specific usage and fine-tuning instructions associated with the paper.

facebookresearch/fairseq
Open GitHub

Official repository for the BeaverTails dataset family, dataset cards, QA-moderation training/evaluation code, and links to released model/data artifacts.

PKU-Alignment/beavertails dataset
Open GitHub

Repository for General AgentBench, the unified benchmark and evaluation framework for general LLM agents.

cxcscmu/General-AgentBench benchmark
Open GitHub

Official repository for the WorfBench benchmark, WorfEval evaluation implementation, data, and experiment code.

zjunlp/WorfBench dataset
Open GitHub

Repository containing code for evaluating MT-BaxBench and MT-SECCODEPLT, the two splits of the MT-Sec evaluation kit.

JARVVVIS/mt-sec benchmark
Open GitHub

OpenAI Codex CLI, cited as a lightweight coding agent running in a terminal and evaluated as one of the agent scaffolds.

openai/codex
Open GitHub

Chapyter is one of the data science agent frameworks benchmarked by DSEval.

chapyter/chapyter
Open GitHub

Jupyter-AI is one of the agent frameworks benchmarked by DSEval.

jupyterlab/jupyter-ai
Open GitHub

Official repository for DSEval, including the evaluation toolkit, benchmark data, results, and scripts.

MetaCopilot/dseval dataset
Open GitHub

Open-source Code Interpreter API implementation directly evaluated as a data science agent framework.

shroominic/codeinterpreter-api
Open GitHub

rllab, the framework released with the paper containing the continuous-control tasks and reference implementations of the evaluated reinforcement-learning algorithms.

rllab/rllab benchmark
Open GitHub

Official repository for the paper's ECG foundation-model benchmark, including evaluation code, preprocessing/configuration assets, and the ECG-CPC framework with checkpoint links.

AI4HealthUOL/ecg-fm-benchmarking benchmark
Open GitHub

GitHub repository for the FaithJudge leaderboard and associated faithfulness evaluation resources.

vectara/FaithJudge benchmark
Open GitHub

Repository containing the modular benchmarking framework, configurations, scripts, results, and visualization code for comparing retrieval strategies in biomedical RAG.

deviprasadbal/RAGHealthcareRetrievalStrategies benchmark
Open GitHub

Repository for MedRag, the toolkit implementing the corpora, retrievers, and LLM configurations used in the study.

Teddy-XiongGZ/MedRAG
Open GitHub

Repository for Mirage, the medical RAG evaluation benchmark introduced in the study.

Teddy-XiongGZ/MIRAGE benchmark
Open GitHub

Repository containing corruption scripts, model-evaluation code, sample execution materials, and links to reduced, corrupted, and verified DUDE and MPDocVQA benchmark data.

DavideNapolitano/VRD-UQA dataset
Open GitHub

Repository containing the agent scaffold, behaviour-analysis pipeline, SFT recipes, scripts, utilities, and evaluation suite associated with Behavior Priming for agentic search.

cxcscmu/Behavior-Priming-for-Agentic-Search system
Open GitHub

Repository containing the authors' BERT-of-Theseus implementation, GLUE scripts, compression examples, and links/instructions for the released MNLI-compressed model.

JetRunner/BERT-of-Theseus
Open GitHub

Repository released by the authors with TensorFlow code for BERT, pretrained BERT_BASE and BERT_LARGE checkpoints, and scripts for replicating major fine-tuning experiments.

google-research/bert
Open GitHub

Repository for reproducing and extending the paper's generator-versus-classifier experiments; it includes the main experiment script, a shell runner, and data files containing synthetic and human-labelled examples.

kinit-sk/multilingual-classifiers-not-generators
Open GitHub

Official implementation and pretrained sample model for the paper's learned reference-free summary reward, including metric comparison and reward-training scripts.

yg211/summary-reward-no-reference
Open GitHub

Official LEAP implementation containing the Python package, prompts, configuration files, data-processing scripts, SFT and DPO training scripts, and evaluation workflows for ALFWorld and WebShop, with InterCode-related code also included in the repository structure.

sanjibanc/leap_llm
Open GitHub

Repository by genai-analytics containing a beyond-black-box-benchmarking folder with benchmark, core, examples, and sdk subfolders.

genai-analytics/publications dataset
Open GitHub

ParlAI crowdsourcing code associated with the paper's multi-session model-chat human evaluation.

facebookresearch/ParlAI
Open GitHub

ParlAI implementation for loading and evaluating the Multi-Session Chat and PersonaSummary tasks introduced by the paper.

facebookresearch/ParlAI dataset
Open GitHub

Code repository identified by the paper for implementing the CRS reliability-evaluation framework.

rohitsalla/CRS
Open GitHub

Official ICML 2023 implementation of DAC, including feature extraction, KNN-score computation, calibration code, experiment scripts, checkpoints, and a CIFAR-10/ResNet18 reproduction workflow.

futakw/DensityAwareCalibration
Open GitHub

Official codebase containing 1dCA data generation, evaluated model implementations, ACT variants, training scripts, evaluation pipelines, and experiment configurations.

RodkinIvan/associative-recurrent-memory-transformer benchmark
Open GitHub

The LM Evaluation Harness supplies the raw, token-normalized, and character-normalized full-sequence scoring methods evaluated by the paper.

EleutherAI/lm-evaluation-harness
Open GitHub

OpenCompass is one of the two evaluation frameworks whose multiple-choice scoring implementation is analyzed and empirically compared.

open-compass/opencompass
Open GitHub

Code repository for Multi-Objective Direct Preference Optimization and its experimental workflow.

ZHZisZZ/modpo implementation
Open GitHub

Repository backing the Agentic Factor Investing project homepage, containing a README, project-framework image, interactive site, and chart data; it is a showcase rather than a complete research-code release.

allenh16/agentic-factor-investing framework
Open GitHub

GoodAI baseline LTM system using a vector database and JSON scratchpad to augment an LLM controller.

GoodAI/goodai-ltm system
Open GitHub

Versioned branch of the GoodAI LTM Benchmark repository containing code, test definitions, experiments, result data, and reports corresponding to the paper.

GoodAI/goodai-ltm-benchmark dataset
Open GitHub

Public repository containing RAG, GRAG, AGRAG, and HRAG prompts; text-to-Cypher and Cypher-to-QA prompts; few-shot query examples; guided-question generation material; the LLM-as-a-Judge rubric; the annual CTI report used for guided questions; and experiment logs.

ait-cti/beyond-vanilla-rag benchmark
Open GitHub

Code for f-DPO with reverse KL, forward KL, Jensen-Shannon, and alpha-divergence regularization, including scripts for IMDB, Anthropic HH, MT-Bench, PPO comparisons, and calibration experiments.

alecwangcq/f-divergence-dpo implementation
Open GitHub

Repository containing the code associated with the Strategic Tactical Agent Reasoning benchmark and its execution framework.

star-nexus/star
Open GitHub

Repository containing the paper's experiments and results for the Agent Assessment Framework.

sa4s-serc/asf benchmark
Open GitHub

Repository containing BIG-bench task definitions, JSON and programmatic evaluation infrastructure, documentation, model score files, BIG-bench Lite resources, and contribution workflows.

google/BIG-bench dataset
Open GitHub

Repository for the paper's supplemental materials, including the coding schema, all repository metadata, raw graph data, and the LLM coding prompt.

ShaokangJiang/cursorrule-supp dataset
Open GitHub

TokenScope is used to extract probabilities for the first decision token where the judge commits to A or B.

Amirresm/tokenscope package
Open GitHub

Repository for the semantic-uncertainty method used as a hallucination-detection baseline.

lorenzkuhn/semantic_uncertainty
Open GitHub

One of the two chest-X-ray datasets combined to construct the paper's four-class federated-learning evaluation corpus.

agchung/Figure1-COVID-chestxray-dataset dataset
Open GitHub

Official BookSum repository containing data-preparation, collection, cleaning, chapterization-support, and paragraph-alignment scripts associated with the paper.

salesforce/booksum dataset
Open GitHub

Repository containing the MuseD evaluation code and associated research artifact.

zhangyipin/mused dataset
Open GitHub

PyTorch implementation of BoostMIS with active-learning components, configuration files, training code, and released MESCC ResNet-50 feature and label artifacts.

wannature/BoostMIS
Open GitHub

Official PyTorch implementation of BOSS for simulated ALFRED experiments associated with the paper.

clvrai/boss framework
Open GitHub

Repository for the paper containing the agent and AI2THOR environment code, experiment entry points, generated production rules and weights, prompts, world-knowledge database, bootstrapping/test logs, experiment spreadsheet, and token-usage figure.

zfy0314/cognitive-agents
Open GitHub

Project repository for the paper's automated peer-review vulnerability evaluation and adversarial robustness study.

Lin-TzuLing/Breaking-the-Reviewer benchmark
Open GitHub

Official repository providing BridgeData V2 data-processing utilities, JAX implementations for several paper methods, training scripts, evaluation scripts, and links to checkpoints and robot setup resources.

rail-berkeley/bridge_data_v2 dataset
Open GitHub

Repository containing preprocessing, KG demos, lab-informed pretraining, diagnosis and heart-failure training scripts, metrics, and notebooks for the DuaLK framework.

humphreyhuu/DuaLK
Open GitHub

Implementation used to generate the saliency-map explanations that formed the H2 baseline condition.

greydanus/visualize_atari
Open GitHub

The gym-sokoban implementation that the authors modified into Sokoban-switch and Sokoban-cells variants for precondition and cost-explanation experiments.

mpSchrader/gym-sokoban
Open GitHub

An aggregated dataset of chess opening names and move sequences used by the paper to create opening-position concept datasets.

lichess-org/chess-openings dataset
Open GitHub

Companion agent layer containing Brain Researcher skills, agent templates, MCP adapters, and AutoResearch evaluation rubrics.

brain-researcher/brain-researcher-agent-kit
Open GitHub

Public Brain Researcher implementation including the Python package, CLI, agent runtime, MCP server with versioned tool contracts, orchestrator, web UI, and deployment recipes.

brain-researcher/brain-researcher-public
Open GitHub

Analysis code, specification ledger, per-hypothesis outputs, run report, and figures for the NeuroMark schizophrenia collaborator case evaluated in the paper.

XinhuiLi/BR-NeuroMark
Open GitHub

Source-code repository for the BCDA study and its algorithmic experiments.

ShironT/bcda implementation
Open GitHub

Code for fitting and extrapolating BNSLs, reproducing the 4-digit-addition experiments and noiseless simulation in Figure 5, and decomposing BNSL into the power-law segments shown in Figure 1.

ethancaballero/broken_neural_scaling_laws
Open GitHub

Repository for downloading/decrypting benchmark data, building or downloading indexes, running supported search agents, evaluating end-to-end and retrieval-only results, reproducing experiments, and integrating custom retrievers.

texttron/BrowseComp-Plus benchmark
Open GitHub

Source code for the CoELA framework and the paper's TDW-MAT and C-WAH experimental implementations, including LLM-agent code and experiment scripts.

UMass-Embodied-AGI/CoELA
Open GitHub

OpenDev, the open-source command-line AI coding agent whose architecture, harness, context engineering, tool system, and lessons learned are described in the paper.

opendev-to/opendev framework
Open GitHub

Open-source Python package calibrated-explanations, including code repository, examples, notebooks, and evaluation/regression scripts for reproducing experiments.

Moffran/calibrated_explanations package
Open GitHub

Repository containing CPR, reset-method baselines, continual-learning environments, experiment configurations, plotting scripts, dependency specifications, and reproduction instructions.

LucMc/continual-learning
Open GitHub

Repository containing experiment data, model configurations, prediction scripts, calibration implementations for agreement/entropy/FSD and baselines, evaluation scripts, scores, and visualization utilities.

veronica320/Calibrating-LLMs-with-Consistency implementation
Open GitHub

Repository containing installation instructions, data preparation, experiment scripts, evaluation code, and analysis for the paper's calibration framework.

kkkevinkkkkk/calibration
Open GitHub

Source code accompanying the paper, with scripts for decoding, extracting internal-consistency information, and evaluating self-consistency variants.

zhxieml/internal-consistency implementation
Open GitHub

The open-source CAMEL library introduced by the paper, including agent implementations, prompts, data-generation pipelines, analysis tools, examples, and links to generated datasets.

camel-ai/camel
Open GitHub

Repository designated by the paper for code, training conditions, and experimental run records.

basetenlabs/cortex
Open GitHub

JSON benchmark file for the paper's freelance-style task suite.

reveondivad/certify dataset
Open GitHub

Repository for the Econometrics-Agent/MetricsAI system that automates econometric analysis through an AI-agent workflow and domain-specific econometric tools.

HKU-Business-AI-Center/Econometrics-Agent framework
Open GitHub

Official repository directory containing the paper's code, data, and experimental setup for writing alignment through edits.

salesforce/creativity_eval
Open GitHub

Repository accompanying the paper, containing the ageval evaluator package, agent tools, baseline evaluators, datasets/labels, experiment configurations, notebooks, and a Gradio app for inspecting annotator outputs.

apple/ml-agent-evaluator framework
Open GitHub

Official repository containing 50 author-specific writing prompts and anonymized JSON evaluation data organized by quality versus style, prompting versus fine-tuning, and expert versus lay judges.

tuhinjubcse/GoodWritingBeGenerative dataset
Open GitHub

Repository linked by the paper as the available code for evaluating GPT models on mock CFA exams.

e-cal/gpt-cfa benchmark
Open GitHub

Official implementation of TextGym, evaluated language-agent configurations, EXE, and the benchmark experiments.

mail-ecnu/Text-Gym-Agents
Open GitHub

Repository for the paper containing the agent_trust experiment code, game prompts, non-repeated and repeated experiment results, environment files, and runnable demos.

camel-ai/agent-trust
Open GitHub

Repository containing resources and code for implementing and experimenting with LLMs for vehicle routing problems, including context materials, LLM framework code, oracle algorithms, and verifier code.

Zhehui-Huang/LLM_Routing benchmark
Open GitHub

Repository for RWE-bench, the benchmark and evaluation environment introduced by the paper for testing LLM agents on real-world evidence generation from medical databases.

somewordstoolate/RWE-bench dataset
Open GitHub

Repository containing the FINSABER framework, backtesting code, strategy interfaces, experiment scripts, documentation, and dataset links for reproducing or extending the paper's benchmark.

waylonli/FINSABER dataset
Open GitHub

Official repository containing code, prompts, modules, tests, and post-processing pipeline for generating and evaluating self-generated counterfactual explanations across the paper's datasets and models.

aisoc-lab/llm-sces implementation
Open GitHub

Repository released by the authors for the confidence-elicitation framework and experiments introduced and evaluated in the paper.

MiaoXiong2320/llm-uncertainty
Open GitHub

Repository containing EnterpriseBench's application data, task-generation components, enterprise tools, evaluation code, and interactive demonstrations.

ast-fri/EnterpriseBench
Open GitHub

Public repository containing the code, method files, data directory, processing notebook, and framework assets used for the paper's MADR experiments.

SangyunLee1027/Code-for-Towards-Faithful-Explainable-Fact-Checking-via-Multi-Agent-Debate
Open GitHub

Repository for IllusionReasoning, the paper's real-world visual-illusion benchmark for evaluating LVLM perceptual and reasoning capabilities.

zhaoliangjie55/EMNLP2026_Illusion dataset
Open GitHub

The isbiased library implements the prediction-shortcut reliance diagnostics used to compute relative accuracy drops in the paper.

MIR-MU/isbiased
Open GitHub

Repository released by the authors for the study's dataset, source code, and detailed evaluation results.

Wendy-1222/SLM_Jailbreak dataset
Open GitHub

Repository released by the authors for the CAP construction pipeline and verifiable agent-as-a-judge evaluation framework.

WarriorXu0302/CAP-Bench
Open GitHub

LangChain repository audited for the default path from model-produced actions to tool execution.

langchain-ai/langchain
Open GitHub

LangGraph repository audited for mandatory value authorization before ToolNode invocation.

langchain-ai/langgraph
Open GitHub

Reference implementation of the paper's deterministic fail-closed authorization gate, including framework integrations and proof scripts.

raceksd-source/scopegate-runtime
Open GitHub

LlamaIndex repository audited for central dispatch, schema validation, and per-call authorization behavior.

run-llama/llama_index
Open GitHub

Stripe Agent Toolkit repository audited for client-side authorization of model-supplied payment arguments.

stripe/agent-toolkit
Open GitHub

Repository for CAR-bench, including benchmark implementation, tools, task and evaluation workflow, results analysis, and documentation for evaluating multi-turn tool-using LLM agents under uncertainty.

CAR-bench/car-bench dataset
Open GitHub

Official CARE-LoRA implementation built on a repository-local PEFT fork, with reproduction workflows for T5-Base and Mistral-7B-v0.3 experiments, gradient-similarity diagnostics, and comparison baselines.

fishandyu/CARE-LoRA implementation
Open GitHub

Companion implementation for the SD3-Medium DreamBooth personalization experiments reported for CARE-LoRA.

fishandyu/CARE-LoRA-Diffusion implementation
Open GitHub

Open-source LLM adversarial robustness toolkit used by the paper to generate adversarial suffixes for the generator-stage jailbreak experiment.

IntelLabs/LLMart
Open GitHub

Repository for the paper's cascaded LLM experiments, including implementation details for the framework evaluated in the paper.

fanconic/cascaded-llms system
Open GitHub

Repository accompanying the paper, containing the P-99 materials, CLAUDE.md instructions, generated problem directories, tests, proof files and the maintained lptp-reference.md.

FredMesnard/LPTP-LLM-P99
Open GitHub

BitTern is presented as an open toolkit for low-cost, high-accuracy post-training ternary quantization and 1.58-bit models.

IntelChina-AI/BitTern implementation
Open GitHub

Repository containing the implementation for causal inference via style transfer for OOD generalisation.

nktoan/Causal-Inference-via-Style-Transfer-for-OOD-Generalisation framework
Open GitHub

Official RSS 2023 implementation of Causal MoMa, including Minigrid and iGibson components, causal-inference and data-collection scripts, and policy-training and testing entry points.

JiahengHu/CausalMoMa
Open GitHub

Repository containing the open-source LLMCert-B implementation, model-query utilities, experiment code, and materials used for the BOLD bias-detector human evaluation.

uiuc-focal-lab/LLMCert-B
Open GitHub

Official repository for the ChAda-ViT architecture, training/evaluation code, and model artifacts.

nicoboou/chadavit
Open GitHub

Microsoft repository containing the Python implementation and supporting materials for the plug-and-play CoNLI hallucination detection and reduction framework.

microsoft/CoNLI_hallucination framework
Open GitHub

Official code repository for Chain-of-Models Pre-Training introduced in the paper.

deep-optimization/CoM-PT
Open GitHub

Repository for the SVAMP math word-problem benchmark.

arkilpatel/SVAMP dataset
Open GitHub

Repository for the ASDiv diverse math word-problem dataset.

chaochun/nlu-asdiv-dataset dataset
Open GitHub

Repository for the AQuA algebraic word-problem dataset.

deepmind/AQuA dataset
Open GitHub

Repository for BIG-bench, including the Date Understanding and Sports Understanding tasks.

google/BIG-bench
Open GitHub

BIG-bench task repository for the StrategyQA question-only evaluation setting.

google/BIG-bench dataset
Open GitHub

Repository associated with the CommonsenseQA benchmark.

jonathanherzig/commonsenseqa dataset
Open GitHub

Repository for the GSM8K grade-school math word-problem benchmark.

openai/grade-school-math dataset
Open GitHub

MIT-licensed Chameleon source repository containing the system implementation, optimization modules, models, benchmark artifacts, and deterministic tests.

RohitSwami33/Chameleon-cybersecurity-ml benchmark
Open GitHub

Repository for the ChannelViT architecture, HCS training/evaluation code, configurations, and pretrained models used to reproduce the paper's experiments.

insitro/ChannelViT implementation
Open GitHub

Repository for the Chat2Workflow benchmark and associated workflow-generation/evaluation resources.

zjunlp/Chat2Workflow dataset
Open GitHub

Public GitHub location for ChatCollab code and data used or produced by the paper.

ChatCollab dataset
Open GitHub

MetaGPT repository, representing a prior meta-programming multi-agent framework compared with ChatCollab.

geekan/MetaGPT framework
Open GitHub

ChatDev repository, representing a prior communicative-agent software-development system compared with ChatCollab.

OpenBMB/ChatDev framework
Open GitHub

Repository for the ChatDev framework and the code/data artifact associated with the paper.

OpenBMB/ChatDev dataset
Open GitHub

Repository explicitly linked by the paper for ChatEDA-Bench, examples of EDA-tool instruction data, API documentation, and the corresponding OpenROAD implementation.

wuhy68/ChatEDAv1
Open GitHub

Official implementation and configuration repository for the ChatEval multi-agent referee framework and its evaluation experiments.

chanchimin/ChatEval
Open GitHub

Repository containing general and targeted synthetic datasets, generation examples, BERT and GPT-2 adapter notebooks, and evaluators for StereoSet, CrowS-Pairs, and BiasTestGPT.

Pengrui-Han/SyntheticDebiasing
Open GitHub

GitHub repository identified by the paper as the location where the dataset can be accessed.

ZihanChen1995/ChatGPT-GNN-StockPredict dataset
Open GitHub

The paper states that the MMF-Trans code has been open sourced at this GitHub URL, with data requiring authorized access. The repository page itself returned 404 during extraction, so its contents could not be inspected.

MMF-Trans framework
Open GitHub

Repository released by the authors for Chronos code and model checkpoints.

amazon-science/chronos-forecasting implementation
Open GitHub

ClawNet repository containing the governed multi-agent social network implementation, including core/gateway, server, desktop, macOS client, and setup components.

hkgai-official/ClawNet framework
Open GitHub

Repository for CMAT: A Multi-Agent Collaboration Tuning Framework for Enhancing Small Language Models, including code, configs, data, docs, and README material describing the framework and evaluation setup.

heimy2000/CMAT dataset
Open GitHub

Official code repository for training and testing CMI-MTL on SLAKE, VQA-RAD, and OVQA, including preprocessing scripts and pretrained-model instructions.

BioMedIA-repo/CMI-MTL
Open GitHub

Paper-linked repository for the introduced framework.

kilian-group/KBevo implementation
Open GitHub

Official CochlScene repository containing baseline training/evaluation code, dataset metadata, a saved model, and data-preparation scripts.

cochlearai/cochlscene
Open GitHub

A GitHub repository listed by the paper as an accompanying curated resource for papers on code as agent harness.

YennNing/Awesome-Code-as-Agent-Harness-Papers
Open GitHub

Official Code as Policies directory in the Google Research repository, containing notebooks for LMP examples, HumanEval, RoboCodeGen, code-versus-language reasoning, reactive controllers, and an interactive demo.

google-research/google-research implementation
Open GitHub

Repository containing released benchmark data derived from SWE-CARE, pipeline scripts for filtering, environment building, test generation, agent resolution, and tool evaluation, plus compressed raw experimental outputs.

c-CRAB-Benchmark/dataset dataset
Open GitHub

Repository for the CodeAssistBench benchmark, datasets, prompts, scripts, and evaluation framework for AI coding assistants on real GitHub issues.

amazon-science/CodeAssistBench dataset
Open GitHub

Open-source repository associated with the CodeScout model family and RL recipe for code localization agents.

OpenHands/codescout system
Open GitHub

GitHub repository for CodeTaste, including benchmark infrastructure, agent/evaluation scripts, documentation, and links to benchmark artifacts and precomputed outputs.

logic-star-ai/codetaste dataset
Open GitHub

Repository for Codev-Bench, the developer-centric repository-level code-completion benchmark constructed with Codev-Agent.

LingmaTongyi/Codev-Bench dataset
Open GitHub

Official implementation of CORY, including IMDB training scripts, GSM8K utilities, TRL-based trainers, environment code, and experiment support files.

Harry67Hu/CORY
Open GitHub

The latest arXiv v3 links this dedicated repository for the newer CogAgent-9B-20241220 release; it is a direct successor artifact to the paper's CogAgent system.

THUDM/CogAgent implementation
Open GitHub

The paper explicitly states that the CogAgent model and code are available in the CogVLM repository, which contains CogAgent-18B resources alongside CogVLM.

THUDM/CogVLM implementation
Open GitHub

Official implementation of the CogVis framework, including inference, staged adapter training, prompt profiles, and benchmark configuration.

KotlinWang/CogVis implementation
Open GitHub

Repository titled KotiJaddu/Masters-Project, described on GitHub as code supporting the author's Master's thesis, with Python source and model folders.

KotiJaddu/Masters-Project implementation
Open GitHub

Official implementation repository containing the ComfyBench benchmark resources, ComfyAgent code, workflow data, agent modules, inference scripts, and evaluation scripts.

xxyQwQ/ComfyBench benchmark
Open GitHub

Repository containing the evaluator-integrity criteria, scored framework census, citation ledger, conformance and verification tooling, and supporting material used to reproduce or check claims from the paper.

idilgozel/evaluator-integrity
Open GitHub

Repository for the LLMWorkflowGenerator project, containing the Python Controller implementation, Android/Termux-oriented workflow automation code, sample files, experiment outputs, and setup instructions.

dos-group/LLMWorkflowGenerator system
Open GitHub

Repository for H2O, the Heavy-Hitter KV-cache eviction method selected as the representative static-sparsification approach.

FMInference/H2O
Open GitHub

Official artifact repository for InfiniGen, the recoverable hierarchical KV-cache selection framework evaluated in the paper.

snu-comparch/infinigen
Open GitHub

Official repository for vLLM, the GPU-resident memory-management framework used as the full-cache baseline in the study.

vllm-project/vllm
Open GitHub

Repository provided by the authors for reproducing the LLM, human-comparison, and algorithmic bandit experiments.

zzy620/LLM-exploration-exploitation implementation
Open GitHub

Official code repository implementing the CompeteAI simulation framework, restaurant environment, agent backends, prompt templates, scenes, database integration, and experiment runner.

microsoft/competeai
Open GitHub

Official code repository for the paper, including model runners, interaction-specific evaluation tasks, attention extraction, and attention manipulation scripts for the OmniReason dataset.

DELTA-DoubleWise/OmniReason
Open GitHub

Repository for the Bitcoin Fee Rate Prediction Project, including data, model notebooks, scripts, result outputs, plots, and citation information for arXiv:2502.01029.

majiangqin/bitcoin dataset
Open GitHub

PyTorch implementation for Compressed Context Memory, including the compression-memory method introduced and evaluated in the paper.

snu-mllab/context-memory
Open GitHub

Repository released by the authors for the compression datasets and the data-collection, processing, and evaluation pipeline used to study the relationship between LLM compression efficiency and downstream capability.

hkust-nlp/llm-compression-intelligence
Open GitHub

ScalePlan packages the operator-level analytical cost model, calibration data, memory/MFU prediction, and parallelism search used by MOSAIC.

dmlc/ScalePlan
Open GitHub

Repository reported by the paper as the code for confidence estimation in LLM-based dialogue state tracking.

jennycs0830/Confidence_Score_DST implementation
Open GitHub

Official repository and Python package for CONFLARE, supporting document loading, cleaning, chunking, calibration-set creation or loading, and conformal retrieval-augmented generation.

Mayo-Radiology-Informatics-Lab/conflare framework
Open GitHub

Repository linked by the authors as the public code for experiments in Conformal Prediction as Bayesian Quadrature.

jakesnell/conformal-as-bayes-quad implementation
Open GitHub

Repository linked by the paper for CCA/SWE-Bench-related materials.

facebookresearch/cca-swebench system
Open GitHub

Open-source PyTorch repository from which the authors selected reproducible GitHub issues requiring specialist debugging.

pytorch/pytorch
Open GitHub

Repository explicitly provided by the authors to reproduce the experiments in the paper.

clinicalml/learn-to-defer implementation
Open GitHub

Companion repository for the paper containing evaluation materials, few-shot and constitutional prompts, and model samples used to document the Constitutional AI experiments.

anthropics/ConstitutionalHarmlessnessPaper
Open GitHub

Repository containing the CLL training implementation, weak-signal generation notebook, model utilities, and scripts for synthetic and real-data experiments.

VTCSML/Constrained-Labeling-for-Weakly-Supervised-Learning
Open GitHub

Python repository for generating PortBench portfolio-theory tasks and evaluating LLM portfolio decisions.

noahardyx/PortBench framework
Open GitHub

Repository accompanying the paper with selected Python code for feature construction, metrics, and the model. The README states that the full proprietary backtesting and continuous-futures processing framework and licensed data are not included.

joelowj/mtl-tsmom implementation
Open GitHub

Repository released by the authors for constructing hierarchical NAS spaces based on CFGs and reproducing the BOHNAS experiments.

automl/hierarchical_nas_construction framework
Open GitHub

GitHub repository identified by the authors as containing the standard performance forecaster model for the Allora network.

allora-network/allora-forecaster framework
Open GitHub

Repository directory containing the implementation and resources for the CoT-MAE contextual masked auto-encoder introduced and evaluated in the paper.

caskcsg/ir implementation
Open GitHub

Repository containing the paper's data and codebase, including data collection materials for the user feedback and annotation tasks.

lil-lab/qa-from-hf dataset
Open GitHub

Repository containing the automated AI research environments and idea execution trajectories introduced in Chapter 4.

NoviScl/Automated-AI-Researcher
Open GitHub

Open-source code, model, and data for the s1 sample-efficient reasoning and budget-forcing work used in Chapters 3 and 4.

simplescaling/s1
Open GitHub

Code artifact for the Synthetic Continued Pretraining and EntiGraph experiments in Chapter 2.

ZitongYang/Synthetic_Continued_Pretraining
Open GitHub

Python repository containing slow-tail, V-shape, and combined max-cash modules, execution scripts, data-schema documentation, and reproducible output paths for the empirical cash-overlay study.

shaun19920309/gd-cash-overlay-filters implementation
Open GitHub

Repository containing code, data artifacts, and reproducible empirical materials for the continuous smooth-signal growth-versus-defensive allocation framework.

ZheliXiong/continuous-smooth-signals-growth-tech-defensive-income-allocation implementation
Open GitHub

Official implementation of the paper's TinyBabyLM construction, model training, GOOD/BAD contrastive generation, mixed-corpus retraining, benchmark evaluation, bootstrap analysis, and figure/table reproduction.

janulm/CD-for-Synthetic-Data-Generation
Open GitHub

Author-linked ConVIRT repository containing the CheXpert 8x200 image-image and text-image retrieval evaluation files; the paper identifies this repository as the public location for its model and collected retrieval datasets.

yuhaozhang/convirt
Open GitHub

Repository identified by the paper as the posted model code for reproducing the ConFIRM workflow.

WilliamGazeley/ConFIRM framework
Open GitHub

The open-source repository for the openCHA framework introduced and demonstrated in the paper.

Institute4FutureHealth/CHA
Open GitHub

Generation script used in constructing the synthetic natural-language-to-Cypher evaluation data for the NeDRex KG silver-standard benchmark.

lucia990/T2C-synthetic-data-generation
Open GitHub

Backend implementation of ChatDRex, including API services, documentation, local setup instructions, and tool-evaluation resources.

SimonSuewerUHH/ChatDRexAPI4J
Open GitHub

Frontend implementation for the ChatDRex conversational user interface and interactive biomedical workflow experience.

SimonSuewerUHH/ChatDRexUI
Open GitHub

GitHub repository associated with the paper's cooperative knowledge distillation method.

MLivanos/Cooperative-Knowledge-Distillation framework
Open GitHub

Official CooperBench repository containing the benchmark package, dataset tooling, task-running CLI, evaluation logic, and cooperative/solo/team settings for coding-agent experiments.

cooperbench/CooperBench dataset
Open GitHub

Repository containing the implementation and configuration details for LSNPC experiments.

huangweipeng7/lsnpc implementation
Open GitHub

Repository containing datasets, inference-refining code, CoT-UQ uncertainty integration code, configuration, pipeline scripts, and result-analysis utilities.

ZBox1005/CoT-UQ implementation
Open GitHub

Repository released by the authors for the COTCAgent system, including agent modules, temporal-analysis code, knowledge-base files, patient-data examples, tests, and interface components.

FrankDengAI/COTCAgent implementation
Open GitHub

Repository containing the generated synthetic datasets, evaluation outputs, and additional t-SNE visualizations for the paper.

Mohdkhalil/Repository-supplementary-for-LAK-25-paper--Creating-Artificial-Students-that-Never-Existed
Open GitHub

Repository for the paper containing CLP/CLSP implementation code, MGSM experiment data, request/merge utilities, and metric scripts.

LightChen233/cross-lingual-prompting implementation
Open GitHub

Repository containing the implementation of CryptoMamba, baseline models, data preprocessing, model training, evaluation metrics, and trading simulation scripts.

MShahabSepehri/CryptoMamba benchmark
Open GitHub

Official CSKV repository with calibration, layer-wise fine-tuning, and inference code for compressed KV-cache models.

wln20/CSKV
Open GitHub

RAGQALeaderboard is used as a benchmark environment for evaluating RAG-QA performance across multi-hop, single-hop, and biomedical question-answering tasks.

AQ-MedAI/RagQALeaderboard dataset
Open GitHub

Repository for the CXL-SpecKV memory manager, speculative prefetcher, FPGA cache engine, host library, driver, framework integration, documentation, and tests.

FastLM/CXL-SpecKV
Open GitHub

Source-code repository for DAG-MoE, including the learned structural aggregation module evaluated in the paper.

JiaruiFeng/DAG-MoE system
Open GitHub

IBM Agentics is the agentic AI framework on which DAO-AI is built; the paper uses its ATypes, logical transduction, and scalable workflow concepts.

IBM/Agentics framework
Open GitHub

GitHub repository identified by the paper as containing DAO-GP source code and datasets.

anonymous273800/DAO-GP dataset
Open GitHub

Repository for ImageNet-Captions, the dataset introduced to pair ImageNet images with original Flickr text and enable controlled language-image versus classification experiments.

mlfoundations/imagenet-captions dataset
Open GitHub

Repository implementing the Bloom-filter-based Data Portrait method described in the paper.

ruyimarone/data-portraits implementation
Open GitHub

FastChat v0.2.5 question JSONL containing 80 diverse evaluation queries.

lm-sys/FastChat dataset
Open GitHub

Repository implementing data-local, ensemble-based, LLM-guided NAS for multiclass multimodal time-series classification, including local executor scripts, remote LLM proposer scripts, schemas, and a Flask control interface.

emilhar/arl framework
Open GitHub

Contains the code for coarse plan generation, fine-grained plan evolution, SFT-data construction, orchestrator and tool-model inference, distributed corpus curation, prompts, and links to released orchestrator and NP checkpoints.

GAIR-NLP/DataOrchestra
Open GitHub

Repository containing the official evaluator, baseline implementations, experimental configurations, and documentation for the DataSpace benchmark.

HKUSTDial/DataSpace benchmark
Open GitHub

Repository cited by the paper for NVIDIA<ef><bf><bd>s Comprehensive Verilog Design Problems benchmark, which provides the selected code-generation and code-comprehension tasks used to evaluate SLMs and LLMs.

NVlabs/verilog-eval dataset
Open GitHub

Repository for the DB2-TransF time-series forecasting model introduced and evaluated in the paper.

SteadySurfdom/DB2-TransF benchmark
Open GitHub

Official repository for the DCQA benchmark, dataset overview, and access instructions.

anranwu-richpo/dcqa dataset
Open GitHub

Microsoft's repository for the DeBERTa implementation, pretrained models, and fine-tuning resources associated with the paper.

microsoft/DeBERTa implementation
Open GitHub

Repository released by the authors for the experiments and data supporting the paper's analysis of probability, memorization, and noisy reasoning in CoT prompting.

aksh555/deciphering_cot
Open GitHub

Code repository for the limited teacher supervision decoding method introduced in the paper.

HJ-Ok/DecLimSup implementation
Open GitHub

Repository containing implementations of the tested actor-critic agents and policy parameterizations, experiment configurations, the Backwashing-PID simulator, source plant data, KL utilities, and plotting/reproducibility support.

haseebs/deconstruct-ac
Open GitHub

An Awesome-style repository organizing MMR papers, datasets, benchmarks, and related surveys around the paper's PerceptionAlignmentReasoning framework and AnswerProcessExecutable evaluation hierarchy.

formula12/Awesome-Multimodal-Mathematical-Reasoning-Perception-Alignment-Reasoning
Open GitHub

AllenNLP contains the official PyTorch ELMo module and scalar-mix implementation for computing the paper's contextual representations.

allenai/allennlp
Open GitHub

Official TensorFlow implementation for training the pretrained bidirectional language model and computing ELMo representations introduced by the paper.

allenai/bilm-tf
Open GitHub

Repository containing the data-pipeline code, hedging simulator, actor-critic training workflow, saved configurations and models, notebooks, evaluation utilities, tests, figures, and paper materials for reinforcement-learning-based hedging of SPX and SPY exposures.

tlucius16/deep-hedging-rl framework
Open GitHub

Public repository containing software code and datasets or scripts for reproducing the baseline algorithms and experiments.

andrepugni/ESC benchmark
Open GitHub

Repository stated by the paper to contain the market simulation and experiment code.

EduardoGarrido90/micro_agents implementation
Open GitHub

Repository linked by the paper as containing the publicly available datasets and code used for the DRL-in-finance experiments.

Andrei-T-Neagu/DRL_in_Finance dataset
Open GitHub

Repository containing code associated with the paper's A2C, DDPG, PPO, ensemble-selection, and backtesting workflow.

AI4Finance-Foundation/Deep-Reinforcement-Learning-for-Automated-Stock-Trading-Ensemble-Strategy-ICAIF-2020 implementation
Open GitHub

Official DeepConf repository implementing offline and online confidence-aware reasoning, voting, trace filtering, warmup thresholding, and early stopping on top of LLM serving backends.

facebookresearch/deepconf
Open GitHub

Repository containing source code, datasets, and supplementary materials associated with the multi-horizon NEM electricity price forecasting benchmark.

GaniMosman/Multi-Horizon-EPF-NEM dataset
Open GitHub

The repository contains DeepAries source code, market data folders, checkpoints, model components, experiment code, and instructions for training and inference.

dmis-lab/DeepAries dataset
Open GitHub

Repository for the paper's VRGA method, including Qwen evaluation code, attention-intervention implementation, dataset interfaces, and a reproducible evaluation pipeline.

Ivine11/VRGA
Open GitHub

Repository path containing DeepPlanning benchmark code, travel and shopping domain runners, evaluation utilities, configuration files, and instructions for reproducing benchmark results.

QwenLM/Qwen-Agent dataset
Open GitHub

Official repository for DeepSeekMath model checkpoints, evaluation materials, replication resources, and inference examples.

deepseek-ai/DeepSeek-Math implementation
Open GitHub

Official DeepSeek-MoE repository linked by the paper for the introduced architecture and released DeepSeekMoE 16B research artifact.

deepseek-ai/DeepSeek-MoE implementation
Open GitHub

The paper's released DeepSpeed-Chat implementation, including end-to-end RLHF examples, training stages, APIs, inference support, tests, and documentation.

microsoft/DeepSpeedExamples
Open GitHub

The DeepSpeed repository contains the framework, code, tutorials, and documentation for large-scale model training and inference, including DeepSpeed-MoE components.

microsoft/DeepSpeed framework
Open GitHub

Repository for DeepTraderX, a deep-learning trading agent running in Threaded-BSE.

armandcismaru/DeepTraderX framework
Open GitHub

Threaded Bristol Stock Exchange repository providing the asynchronous market simulator and working trading agents used for DTX training data and experiments.

MichaelRol/Threaded-Bristol-Stock-Exchange benchmark
Open GitHub

Repository identified by the paper as containing code and data for the finance hallucination experiments.

mk322/fin_hallu benchmark
Open GitHub

Software suite described and used by the paper to calculate reward-representable policy orderings, non-trivial simplifications, and reward functions representing them for specified environments and policy sets.

nikihowe/reward-hacking-paper
Open GitHub

Official DeFT research repository containing Triton implementations, tree templates, benchmark scripts, and reproduced performance results for the paper's attention methods.

LINs-lab/DeFT
Open GitHub

Code, documentation, and demos for ToolUniverse, the ecosystem for building AI scientists from language models, reasoning models, and agents.

mims-harvard/ToolUniverse framework
Open GitHub

Official implementation repository released for the Demonstrate-Search-Predict framework.

stanfordnlp/dsp
Open GitHub

Salesforce AI Research repository for FinDAP: Demystifying Domain-adaptive Post-training for Financial LLMs, including framework materials, training scripts, evaluation guidance, and links to model/data artifacts.

SalesforceAIResearch/FinDAP dataset
Open GitHub

Repository associated with the AgentFail dataset and website for failure lifecycle data and analyses.

Jenna-Ma/JaWs-AgentFail dataset
Open GitHub

Repository containing DPR training, retrieval, evaluation, data-processing tools, configurations, and released model resources.

facebookresearch/DPR
Open GitHub

Code repository provided by the paper for DAAM implementations and the IEMOCAP, AG News, CIFAR-100, LoRA-comparison, and appendix experiments.

gioannides/DAAM-PEFT-paper-code
Open GitHub

GitHub repository for the TSCC2019 competition data used as real Hangzhou traffic data in the paper's traffic-light-control benchmark.

tianrang-intelligence/TSCC2019 dataset
Open GitHub

Repository containing the released loss data and analysis scripts/notebooks for optimal hyperparameter estimation, learning-rate and batch-size scaling, Chinchilla/Skaling fitting, compute-optimal scaling, and uncertainty estimation.

OpenEuroLLM/dense_english_scaling_laws
Open GitHub

Official code repository for the DEPS/MC-Planner implementation introduced and evaluated in the paper.

CraftJarvis/MC-Planner
Open GitHub

Repository containing the blank Design-OS template, design-case artifacts, prompt files, simulation code, plots, and verification report used to support reuse and replicability testing.

bankh/design-os framework
Open GitHub

GitHub repository for the TalkTuner paper, including code and data for a dashboard that visualizes and controls a chatbot LLM's internal user model.

yc015/TalkTuner-chatbot-llm-dashboard dataset
Open GitHub

GitHub repository for the evolutionary multi-objective neural architecture search approach introduced in the paper.

DevilYangS/EMO-NAS-CD framework
Open GitHub

Repository linked by the paper for reproducing or inspecting the OAS generation experiment.

marques-vinicius/OAGen implementation
Open GitHub

Repository containing code, model files, feature data, and training scripts for the paper's physics-guided bearing prognostic and uncertainty-quantification experiments.

itxwaleedrazzaq/uqpcnn_rul
Open GitHub

Repository describing DeXposure-FM as a time-series graph foundation model for forecasting inter-protocol credit exposure, with scripts for experiments, macroprudential tools, checkpoints, and dataset helpers.

EVIEHub/DeXposure-FM framework
Open GitHub

The code URL reported in the paper's v2 HTML/PDF abstract for the DeXposure-FM project.

EVIEHub/graph-dexposure framework
Open GitHub

Official MuJoCo Allegro Hand environments used for the DIME simulation experiments.

NYU-robot-learning/dime_env
Open GitHub

Official scripts for collecting and processing demonstrations for the rotation, spinning, and flipping tasks.

NYU-robot-learning/DIME-Demonstrations
Open GitHub

Official package for inverse-kinematics-based teleoperation of an Allegro Hand attached to a Kinova Arm.

NYU-robot-learning/DIME-IK-TeleOp
Open GitHub

Official implementation repository containing the nearest-neighbor imitation component and links to the other DIME packages.

NYU-robot-learning/DIME-Models
Open GitHub

Official code repository for DiaLLM, including supervised fine-tuning, PPO, reward-modeling, evaluation scripts, data folders, and result folders.

WeijieyingRen/DiaLLMs
Open GitHub

Code repository for the DiffLOB regime-conditioned diffusion model.

ZhuoHan1998/DiffLOB implementation
Open GitHub

Autonomous-driving decision-making simulator used for DiLu's closed-loop experiments and domain-shift evaluations.

eleurent/highway-env
Open GitHub

Code repository for running and visualizing the DiLu closed-loop LLM driving framework.

PJLab-ADG/DiLu implementation
Open GitHub

Official CitySim vehicle-trajectory dataset repository used as the source of one transferred DiLu memory module.

UCF-SST-Lab/UCF-SST-CitySim1-Dataset
Open GitHub

Training code for supervised fine-tuning followed by DPO preference learning on causal Hugging Face language models, with dataset and trainer utilities.

eric-mitchell/direct-preference-optimization implementation
Open GitHub

Google Java Format repository; the paper uses Newlines.java from this repository for generated-vs-original unit-test comparison.

google/google-java-format
Open GitHub

Junit5 Modular World sample module; the paper uses a Flavor.java code snippet from this repository for prompt-based test generation and comparison.

junit-team/junit5-samples
Open GitHub

Contains code, data, configurations, and scripts for the FineLogic training and evaluation experiments.

YujunZhou/FineLogic benchmark
Open GitHub

Author repository for the GSM8K-AI-SubQ reasoning dataset and baselines for distilling LLM decomposition abilities into compact language models.

DT6A/GSM8K-AI-SubQ dataset
Open GitHub

Official source code repository for the Distilling step-by-step method introduced in the paper.

google-research/distilling-step-by-step implementation
Open GitHub

Repository containing code for the supervised and unsupervised LLM uncertainty experiments reported in the paper.

gahdritz/llm_uncertainty implementation
Open GitHub

GitHub repository released by the authors for Distributed Conformal Prediction via Message Passing.

HaifengWen/Distributed-Conformal-Prediction implementation
Open GitHub

Official code repository for the Division-of-Thoughts framework introduced and evaluated in the paper.

tsinghua-fib-lab/DoT framework
Open GitHub

Official repository for DNABERT-2 code, model usage, pre-training/fine-tuning scripts, and GUE benchmark resources.

MAGICS-LAB/DNABERT_2 implementation
Open GitHub

Official self-contained SayCan implementation in a simulated tabletop environment using a UR5 setup, ViLD affordances, GPT-3 planning, and a CLIPort pick-and-place policy.

google-research/google-research implementation
Open GitHub

Official SayCan dataset files mapping natural-language user instructions and initial conditions to possible solution plans.

say-can/say-can.github.io dataset
Open GitHub

Repository containing code and notebooks associated with the S&P 500 graph-neural-network forecasting project.

waderylan/sp500-gnn implementation
Open GitHub

The paper-associated repository curates video-LLM systems, visual encoders, datasets, code links, figures, and citation information related to the review.

Darcyddx/Video-LLM
Open GitHub

Repository linked by the paper for code and the SelfAware research artifact used to evaluate whether LLMs recognize unanswerable questions.

yinzhangyue/SelfAware dataset
Open GitHub

Repository containing code and data for evaluating uncertainty estimation in LLM instruction-following.

apple/ml-uncertainty-llms-instruction-following dataset
Open GitHub

Repository for Stanford Alpaca, whose Alpaca-7B model is directly evaluated in the paper's altered task-definition experiments.

tatsu-lab/stanford_alpaca
Open GitHub

Public source-code repository for the multimodal repair-recognition system introduced and evaluated in the paper.

haanh764/multimodal_repair_recognition
Open GitHub

Contains VLN-HAMT and VLN-DUET integrations, training/inference scripts, requirements, and links to R2R-Imagine generations, metadata, extracted features, and trained HAMT-Imagine/DUET-Imagine checkpoints.

akhilperincherry/VLN-Imagine
Open GitHub

Repository reported by the paper as containing the prediction model implementation and experiment data.

majuanjuan/Doexpertsperformbetter dataset
Open GitHub

Repository for the DocAgent multi-agent code documentation generation framework introduced by the paper.

facebookresearch/DocAgent framework
Open GitHub

Repository accompanying the paper, linking to hosted C4 data and providing a public discussion venue for documenting additional corpus issues.

allenai/c4-documentation
Open GitHub

Repository for the code, human-subject study materials, results, and supplementary materials associated with the Persona framework and AAAI 2025 paper.

YODA-Lab/Persona framework
Open GitHub

Semantic routing package for routing inputs by embedding or intent similarity.

aurelio-labs/semantic-router package
Open GitHub

Framework using repeated generations, verification prompts, and confidence estimates to decide whether to escalate to larger models.

automix-llm/automix framework
Open GitHub

AWS multi-agent orchestration framework that includes prompt-based routing or agent selection patterns.

awslabs/multi-agent-orchestrator framework
Open GitHub

Implementation associated with routing prompts to pre-trained experts after fine-tuned meta-model categorisation.

godcherry/ExpertTokenRouting implementation
Open GitHub

Implementation associated with deciding whether a query requires a complex prompting strategy.

imagination-research/sot implementation
Open GitHub

Iterative multi-agent code generation system using execution success as a routing signal.

JieyuZ2/EcoAssistant system
Open GitHub

LLM routing implementation associated with assessing model adequacy through multiple responses and ground-truth comparison.

kvadityasrivatsa/llm-routing implementation
Open GitHub

Routing-agent implementation using synthetic data and small classifiers for classification-based routing.

lamini-ai/llm-routing-agent implementation
Open GitHub

Orchestrator implementation using decoder-only LLM representations for routing or model selection.

Leeroo-AI/leeroo_orchestrator system
Open GitHub

Framework for serving and evaluating routers that choose between LLMs using preference-oriented routing strategies.

lm-sys/RouteLLM framework
Open GitHub

Task-planning framework in which an LLM selects among models or tools based on descriptions and user tasks.

microsoft/JARVIS framework
Open GitHub

Implementation assessing consistency across reasoning representations for cascade-style routing.

MurongYue/LLM_MoT_cascade implementation
Open GitHub

OpenAI multi-agent orchestration framework discussed as an example of prompt-based routing practice.

openai/swarm framework
Open GitHub

Fine-tuned model framework for API call generation, discussed as treating routing as a code generation problem.

ShishirPatil/gorilla implementation
Open GitHub

Framework for reducing LLM application cost using LLM cascades and related strategies.

stanford-futuredata/Frugalgpt framework
Open GitHub

Adaptive RAG framework that routes among no retrieval, single-step retrieval, and multi-step retrieval paths according to query complexity.

starsuzi/Adaptive-RAG framework
Open GitHub

Code and data for a multi-LLM routing benchmark and evaluation framework.

withmartian/routerbench dataset
Open GitHub

GitHub repository containing the Indonesian financial-domain language-model code and post-trained IndoBERT models.

intanq/indonesian-financial-domain-lm implementation
Open GitHub

Implementations of original, sparse, continuous, and related statistical jump models.

Yizhan-Oliver-Shu/jump-models framework
Open GitHub

Code and model artifacts for Reinforced Token Optimization, the paper's DPO-derived token-reward and PPO/RL alignment method.

zkshan2002/RTO implementation
Open GitHub

Official RSS 2024 repository implementing DrEureka reward generation, reward-aware physics priors, domain-randomization generation, and the forward-locomotion and globe-walking environments.

eureka-research/DrEureka
Open GitHub

Provides the Go1 soccer simulation environment, PPO training scripts, pretrained policy evaluation, configuration, and documentation accompanying the paper.

Improbable-AI/dribblebot
Open GitHub

Official repository for the paper's LLM-based autonomous-driving demonstrations, including closed-loop HighwayEnv code, scenario assets, and common-sense reasoning materials.

PJLab-ADG/DriveLikeAHuman
Open GitHub

Official repository containing the DS-1000 data, execution-based evaluators, environment files, inference scripts, and released baseline results.

xlang-ai/DS-1000 dataset
Open GitHub

Repository for evaluating LLMs on the DSBC dataset, including response generation, LLM-as-judge evaluation, command-line usage, and dataset evaluation utilities.

traversaal-ai/DSBC-Data-Science-Task-Evaluation dataset
Open GitHub

Repository for DSPy, the programming model and compiler introduced by the paper.

stanfordnlp/dspy framework
Open GitHub

Code repository for the DSTCGCN traffic forecasting model introduced and evaluated in the paper.

water-wbq/DSTCGCN implementation
Open GitHub

Official D-DiT repository with environment setup, pretrained 512-base and supervised-fine-tuned checkpoints, inference examples, data formats, training configurations, and distributed training commands.

zijieli-Jlee/Dual-Diffusion system
Open GitHub

A scikit-learn-style implementation of a collection of statistical jump models, including the model family used to identify asset-specific market regimes.

Yizhan-Oliver-Shu/jump-models package
Open GitHub

Repository indicated by the paper as the code website for DVGNN.

gorgen2020/DVGNN implementation
Open GitHub

Repository containing experimental analysis, tools, and code associated with dynamic design of machine-learning pipelines via metalearning.

ealcobaca/dynamic-design-machine-learning-pipelines framework
Open GitHub

The first author's repository providing implementations of statistical jump models, including methods relevant to the sparse jump-model analysis used in the paper.

Yizhan-Oliver-Shu/jump-models package
Open GitHub

Official GitHub repository containing code for the paper Dynamic Graph Convolutional Network with Attention Fusion for Traffic Flow Prediction.

trainingl/AFDGCN implementation
Open GitHub

Repository for the Dynamic Meta-Learning for Adaptive XGBoost-Neural Ensembles implementation, including source code and test datasets as described by the paper.

aasedek/Adaptive-XGBoost-Neural-Network-Ensemble framework
Open GitHub

Repository containing data and code for the ICLR 2023 PromptPG paper, including the TabMWP dataset and implementations/evaluation scripts for GPT-3, UnifiedQA, TAPEX, and PromptPG experiments.

lupantech/PromptPG dataset
Open GitHub

Python repository containing pseudo-answer generation, silver-label construction, input processing, multi-task training, inference, and evaluation scripts for the six reported datasets.

XieZilongAI/E2E-AFG framework
Open GitHub

Repository containing the EasyRAG pipeline, ingestion code, retrievers, rerankers, prompt templates, challenge scripts, Docker deployment, FastAPI service, Streamlit WebUI, and processed challenge assets.

BUAADreamer/EasyRAG dataset
Open GitHub

The EasyTool directory in Microsoft JARVIS contains implementation code, requirements, generated tool-instruction data for ToolBench and RestBench, FuncQA data, and commands for reproducing the reported evaluations.

microsoft/JARVIS implementation
Open GitHub

Official Microsoft Research implementation of the Sui Generis scoring pipeline introduced in the paper.

microsoft/SuiGeneris
Open GitHub

Public GitHub repository for the study's replication materials, mirrored in Zenodo.

codescene-research/echoes-of-ai-emse-2025 dataset
Open GitHub

Official code repository for the EconAgent paper, including simulation scripts, configuration, data, and AI-Economist foundation components.

tsinghua-fib-lab/ACL24-EconAgent
Open GitHub

Repository containing the EDA augmentation implementation and code associated with the paper's experiments.

jasonwei20/eda_nlp
Open GitHub

Code repository linked by the paper for the ChatSim system and its introduced simulation pipeline.

yifanlu0227/ChatSim implementation
Open GitHub

Code repository for CAID, including the multi-agent workflow where a central manager delegates tasks to engineer agents that run asynchronously in isolated git worktrees, plus scripts and task modules for Commit0 and PaperBench experiments.

JiayiGeng/CAID system
Open GitHub

NovGrid extends MiniGrid with a generalized novelty generator so environment properties and dynamics can change and agents can be evaluated on adaptation to those changes.

eilab-gt/NovGrid dataset
Open GitHub

OWL is compared against Efficient Agents in the agent framework evaluation.

camel-ai/owl framework
Open GitHub

Smolagents is compared against Efficient Agents and OWL in the agent framework evaluation.

huggingface/smolagents framework
Open GitHub

Repository linked by the paper as its code resource, associated with the OAgents/Efficient Agents work.

OPPO-PersonalAI/OAgents framework
Open GitHub

Public code repository for the DASH architecture-search algorithm introduced and evaluated in the paper.

sjunhongshen/DASH framework
Open GitHub

Official repository for the NAACL 2021 paper, containing data and model resources needed to reproduce reported results.

luyang-huang96/LongDocSum implementation
Open GitHub

fairseq example code and resources for training and evaluating the paper's MoE language models.

pytorch/fairseq implementation
Open GitHub

Source-code repository for Helium, the workflow-aware LLM serving system proposed and evaluated in the paper.

mlsys-io/helium_demo framework
Open GitHub

Public source repository for vLLM, the LLM serving system introduced and evaluated in the paper.

vllm-project/vllm
Open GitHub

The paper's vLLM fork and branch containing its Activated LoRA serving and cross-model prefix-cache reuse implementation.

tdoublep/vllm implementation
Open GitHub

Maintained companion repository for the survey's literature organization and updates on efficient multimodal large language models.

lijiannuist/Efficient-Multimodal-LLMs-Survey
Open GitHub

Repository for the Temporal Neural Common Neighbor model introduced in the paper.

GraphPKU/TNCN implementation
Open GitHub

Repository containing SDPG agents, configurations, Genesis environments, models, training/evaluation scripts, baseline dependencies, and Go2 hardware-related files.

HaoxiangYou/SDPG
Open GitHub

Official repository for ActPRM, containing the active-PRM package, training and ProcessBench evaluation scripts, examples, and links to released models and data.

sail-sg/ActivePRM
Open GitHub

Official experiment code for the paper's expert-wise mixed-precision MoE quantization method.

nowazrabbani/moe_quantization
Open GitHub

Repository released by the authors for PTA-IRT code and data.

DeepSoftwareAnalytics/PTA-IRT
Open GitHub

Repository associated with EfficientLLM for the benchmark codebase, metric API, evaluation workflows, and planned model artifacts.

DLYuanGod/EfficientLLM
Open GitHub

Repository containing the code, experiment configurations, and model artifacts for the structured state-space work discussed in the paper.

HazyResearch/state-spaces
Open GitHub

Official code repository for generating embeddings and pseudo-labels, collecting training dynamics, computing importance scores, selecting ELFS coresets, and training CIFAR and ImageNet classifiers.

eltsai/elfs
Open GitHub

Official repository containing scripts and links to reconstruct the ELI5 dataset, build support documents and splits, format multi-task data, train and evaluate the reported models, and use the released pretrained checkpoint.

facebookresearch/ELI5 dataset
Open GitHub

Repository implementing the paper's ELSAA hybrid attention mechanism, including sortLSH sparse attention, RACE low-rank attention, denominator-aware fusion, causal variants, and experiment scripts.

mahdiheidari721/ELSAA
Open GitHub

Repository for EmbedLLM materials, described by the paper as containing the dataset, code, and embedder for further research and application.

richardzhuang0412/EmbedLLM dataset
Open GitHub

Official repository for EMMA, containing the LAVIS-based ALFWorld finetuning setup, links to released training artifacts, and trajectory/experiment logs for EMMA and the Reflexion LLM teacher.

stevenyangyj/Emma-Alfworld
Open GitHub

Repository for the Embodied Web Agents project, including web environment hosting instructions and model-running folders for indoor, outdoor, and geolocation tasks.

Embodied-Web-Agent/Embodied-Web-Agent system
Open GitHub

Python SDK that lets users obtain drone observations and issue control actions through the Embodied City online API.

tsinghua-fib-lab/embodied-city-python-sdk package
Open GitHub

Repository containing simulator-related materials, datasets, task code, prompts, VLN code, and documentation for the EmbodiedCity benchmark.

tsinghua-fib-lab/EmbodiedCity dataset
Open GitHub

Official code repository for the EMERGE framework, including data processing, retrieval, fusion, training, and evaluation components.

yhzhu99/EMERGE
Open GitHub

Code and experiment resources for measuring how context-characteristic sensitivity changes across instruction-fine-tuning stages.

copenlu/context-characteristics-sensitivity implementation
Open GitHub

Repository named Emoji-Embedding-For-Finance, cited throughout the paper as the source for model-vs-BERT comparisons, emoji frequencies, BTC/VCRIX figures, sentiment time series, and trading-strategy outputs.

QuantLet/Emoji-Embedding-For-Finance implementation
Open GitHub

Repository linked as 'Code' from the paper's official project page; its README states that it maintains an overview/implementation of the Habitat-MAS benchmark and EMOS multi-agent system and contains a habitat-mas component plus instructions for running the EMOS demo.

SgtVincent/EMOS benchmark
Open GitHub

Official repository for the paper, containing code and data artifacts for generating analysis-report features and training the hybrid asset pricing model.

chengjunyan1/AAPM implementation
Open GitHub

Core open-source Academy implementation for building and deploying stateful actors and autonomous agents across distributed and federated research infrastructure.

academy-agents/academy
Open GitHub

Code for generating and scoring multiple-choice questions from scientific papers, used as the implementation basis for the paper's information-extraction case study.

auroraGPT-ANL/MCQ-and-SFT-code
Open GitHub

Repository released by the authors for ILM code, data, trained models, and human-evaluation outputs/responses.

chrisdonahue/ilm implementation
Open GitHub

Official code and data repository for the ALCE benchmark and evaluation framework introduced in the paper.

princeton-nlp/ALCE dataset
Open GitHub

Pre-release Python/PyTorch code for building chart-image datasets, splitting data, training the ResNet trader, and inferring triple-I weights.

ZhuZhouFan/TWMA implementation
Open GitHub

Repository for implementing Iter-CoT, the paper's iterative bootstrapping method for chain-of-thought prompting.

GasolSun36/Iter-CoT implementation
Open GitHub

Preferred Multi-turn Benchmark for Finance in Japanese, used by the paper to evaluate generation quality across financial dialogue tasks.

pfnet-research/pfmt-bench-fin-ja benchmark
Open GitHub

Pathway is used as the vector store implementation in the paper's retrieval system.

pathwaycom/pathway implementation
Open GitHub

Stated repository for the modular Python prototype implementing the neuro-symbolic ontology-based LLM validation pipeline.

ruslanmv/Neuro-symbolic-interaction system
Open GitHub

Repository containing implementation for the PSX interpretability algorithms and models discussed in the paper.

sahar-arshad/PSX-Interpretability implementation
Open GitHub

Repository indicated by the paper as containing code for AdvDistill; the URL returned 404 when fetched during extraction.

shreyansh-2003/AdvDistill implementation
Open GitHub

Official CASE inference-code repository linked directly by the paper. The repository implements the context-aware semantic embedding pipeline with a scikit-learn-style transformer interface.

SAP-samples/case implementation
Open GitHub

Public implementation of the structured-matrix Transformer enhancement framework introduced in the paper.

newbeezzc/MonarchAttn framework
Open GitHub

Code repository for the Focus reference-free uncertainty-based hallucination detector.

zthang/focus
Open GitHub

Official PyTorch implementation of JADE and distribution point for the text portion of the CC3M-QA-DC dataset.

CASIA-IVA-Lab/OPT_Questioner
Open GitHub

Official repository for the paper containing scripts for question-driven caption generation, GPT-3.5 answer generation, preprocessing and evaluation, plus released captions and prediction files.

ovguyo/captions-in-VQA
Open GitHub

Repository containing the LoT prompting implementation, scripts for CoT and LoT experiments, requirements, and quick-run instructions.

xf-zhao/LoT framework
Open GitHub

Gym-PushT is the simulation environment used for ENPIRE's heuristic-learning autoresearch comparison and is explicitly cloned for the Push-T setup.

huggingface/gym-pusht
Open GitHub

GitHub repository cited by the paper as the source of the second obesity classification dataset.

pymche/Machine-Learning-Obesity-Classification dataset
Open GitHub

Repository reported by the paper as containing code, curated data pointers, generated figures, and result tables for the reproduction and robustness analysis.

ZheliXiong/Ensemble-RL-through-Classifier-Models implementation
Open GitHub

Official repository containing the code used to train and evaluate the ME-DST models introduced in the paper.

WRF32-10/ME-DST
Open GitHub

Implementation of faithfulness-aware decoding used for the advanced-decoding baseline.

amazon-science/faithful-summarization-generation implementation
Open GitHub

FactPEGASUS code and models used to test whether the proposed augmentation transfers to another factuality-aware contrastive pipeline.

meetdavidwan/factpegasus implementation
Open GitHub

Official CLIFF implementation used as the contrastive-learning baseline and as the training framework combined with the proposed counterfactual augmentation.

ShuyangCao/cliff_summ implementation
Open GitHub

Official ERA repository containing the EPL training code, online RL framework, environment setup, evaluation integration, and links to the curated prior datasets.

Embodied-Reasoning-Agent/Embodied-Reasoning-Agent
Open GitHub

Repository hosting the ESG-FTSE corpus of news articles with ESG relevance labels.

mariavpavlova/ESG-FTSE-Corpus dataset
Open GitHub

Repository released by the authors containing code for the agentic benchmark assessment and related experiments.

uiuc-kang-lab/agentic-benchmarks benchmark
Open GitHub

Repository associated with the paper's contamination-detection work for LLM evaluation. The paper links it as code and data; the currently visible README describes a lightweight tool for identifying and analysing potential contamination without access to LLM training data.

liyucheng09/Contamination_Detector dataset
Open GitHub

Repository URL listed by the paper for the implementation of the confidence IQN experiments.

YHL04/confidenceiqn implementation
Open GitHub

Repository containing the paper's data-processing, rephrasing, fine-tuning, detector-evaluation, and post-processing code for reproducing the reported experiments.

eth-sri/malicious-contamination
Open GitHub

A production-grade evaluation toolkit for LLM agent outputs, including metrics for diversity, reliability, cascade uncertainty, perturbation consistency, consistency, factual grounding, hallucination, explainability, and drift.

mukund1985/llm-eval-toolkit framework
Open GitHub

Repository containing the AGENTbench harness used to evaluate coding agents under NONE, LLM, and HUMAN repository-level context settings on AGENTbench and SWE-bench Lite.

eth-sri/agentbench benchmark
Open GitHub

Repository associated with the CogEval cognitive-map study containing conversation artifacts, task templates, and a conversation visualizer.

cogeval/cogmaps
Open GitHub

Public GitHub repository released by the authors to support reproducibility and further research on financial relationship graph evaluation.

FreddieNIU/Financial-Graph-Evaluation dataset
Open GitHub

Repository containing the HumanEval hand-written programming benchmark and code for evaluating generated solutions.

openai/human-eval dataset
Open GitHub

Repository described as a resource hub for understanding, detecting, and mitigating biases in financial-domain LLMs, including a Structural Validity Checklist, a literature review dashboard, and an automatic bias detection dashboard.

Eleanorkong/Awesome-Financial-LLM-Bias-Mitigation framework
Open GitHub

Official repository containing data-download scripts, annotation tools, data-construction code, API and open-source model inference code, evaluation code, and reward-model evaluation utilities.

LesterGong/MMRB
Open GitHub

Repository containing the code released for the ClinMM-Bench study.

ruiyang-medinfo/ClinMM
Open GitHub

Repository associated with the M4 competition dataset and methods, used by the paper for benchmark data and base learner forecasts.

Mcompetitions/M4-methods dataset
Open GitHub

Repository containing the source code for the paper's FFORMA and ES-RNN ensemble experiments.

Pieter-Cawood/FFORMA-ESRNN benchmark
Open GitHub

Repository for LogiEval, the paper's prompt-style logical-reasoning benchmark suite for evaluating large language models.

csitfun/LogiEval benchmark
Open GitHub

LogiQA 2.0 repository referenced for the benchmark and for the newly constructed 2022-onward out-of-distribution logical-reasoning data.

csitfun/LogiQA2.0 dataset
Open GitHub

Official repository containing the LoCoMo data release, conversation-generation code, prompts, and evaluation scripts.

snap-research/LoCoMo dataset
Open GitHub

Phoenix is cited among tools that provide analytics and evaluation orchestration capabilities for agent or LLM evaluation.

Arize-ai/phoenix
Open GitHub

DeepEval is listed among tools that support analytics, evaluation orchestration, and debugging for LLM or agent evaluation workflows.

confident-ai/deepeval
Open GitHub

OpenAI Evals is discussed as an open-source framework for specifying evaluation tasks and metrics and automating execution and reporting.

openai/evals
Open GitHub

HAL is cited as a holistic agent leaderboard or harness for centralized and reproducible agent evaluation.

princeton-pli/hal-harness benchmark
Open GitHub

Inspect AI is cited as a framework for large language model evaluations and included among evaluation tooling examples.

UKGovernmentBEIS/inspect_ai
Open GitHub

Repository named in the paper for ICD coding explainability evaluation resources, including RD-IV-10-related artifacts and generated rationales.

mingyangligithub/ICD-Coding-Explainability-Evaluation dataset
Open GitHub

Paper-provided implementation link for the evaluation-awareness scaling-law experiments; the arXiv HTML resolves through an Anonymous Github mirror.

eval-awareness-scaling-laws benchmark
Open GitHub

Official repository for the survey, containing the paper's benchmark reference framework, figures, and benchmark bibliography.

YHPeter/Awesome-RAG-Evaluation
Open GitHub

Repository containing the EviReform source code, configuration, prompt assets, released evaluation data, index artifacts, metric/statistics code, and reproduction instructions.

XrazyMee/EviReform
Open GitHub

R code, datasets, and experiment scripts for EvoAAA, the evolutionary autoencoder architecture search methodology introduced in the paper.

fcharte/EvoAAA dataset
Open GitHub

Repository containing code and technical details for the Multi-Agent Scoring System for essay assessment.

AzizovDilshod/Multi-agent-System-for-Essay-Assessment system
Open GitHub

OpenAI's released distributed implementation of the evolution-strategies algorithm studied in the paper.

openai/evolution-strategies-starter
Open GitHub

Repository providing BanditBench and inference code; the paper also notes installation via pip install banditbench.

allenanie/EVOLvE package
Open GitHub

Official repository for the CodeAct framework, evaluation code, CodeActAgent deployment components, scripts, and links to the released data and models.

xingyaoww/code-act implementation
Open GitHub

Repository identified by the paper as containing the code and data for the software-developing agent framework evaluated in the study.

OpenBMB/ChatDev dataset
Open GitHub

The repository is linked by the paper as the location for data and code and contains files including a notebook, UNSW-NB15 dataset archive, and feature metadata.

pcwhy/XML-IntrusionDetection dataset
Open GitHub

Repository named by the authors as containing all code for the experiments that apply feature importance, SHAP, and LIME to PPO portfolio-management predictions.

aleedelarica/XDRL-for-finance implementation
Open GitHub

Existing Reinforcement learning in portfolio management repository that the authors identify as the starting framework for integrating explainability into a PPO-based model.

deepcrypto/Reinforcement-learning-in-portfolio-management- framework
Open GitHub

The cpath package implements the paper's counterfactual-path method for R and Python.

pievos101/cpath package
Open GitHub

UnifiedSKG is the structured knowledge grounding model used in the paper's text-to-SQL case study.

HKUNLP/UnifiedSKG system
Open GitHub

Official implementation of Pattern-Exploiting Training and iterative PET introduced in the paper.

timoschick/pet implementation
Open GitHub

Repository for the x-stance multilingual stance-detection dataset evaluated by the paper.

ZurichNLP/xstance dataset
Open GitHub

Repository containing the benchmark data, task files, representative logs, and evaluation scripts for AutoGen, MetaGPT, and TaskWeaver.

lurf21/Agent_Evaluation_Framework dataset
Open GitHub

Repository for the MachineSoM experiments, including sampled evaluation data, prompts, agent-society simulation code, agent-count/round/strategy ablations, ANOVA evaluation, and conformity/consensus plotting utilities.

zjunlp/MachineSoM
Open GitHub

Implements SimCLR-style HAR training and evaluation on MotionSense, including preprocessing, transformations, model definitions, utilities, and a demonstration notebook.

iantangc/ContrastiveLearningHAR
Open GitHub

GPT-Engineer is listed as a code-domain LLM-based single-agent system and cited as a GitHub project that generates code repositories from prompts.

AntonOsika/gpt-engineer system
Open GitHub

GPTresearcher is listed as a research-domain LLM-based autonomous agent for online comprehensive research.

assafelovic/gpt-researcher system
Open GitHub

AIlegion is listed as a universal LLM-powered autonomous agent platform.

eumemic/ai-legion implementation
Open GitHub

LoopGPT is listed as a universal modular Auto-GPT-style LLM agent framework.

farizrahman4u/loopgpt framework
Open GitHub

AGiXT is listed as a universal AI automation platform with instruction management, memory, and plugins.

Josh-XT/AGiXT
Open GitHub

LangChain is discussed as an open-source framework supporting LLM-based agent software development and tool integration.

langchain-ai/langchain framework
Open GitHub

DemoGPT is listed as a code-support LLM-based agent system for creating LangChain applications through prompts.

melih-unsal/DemoGPT system
Open GitHub

AgentGPT is discussed as an agent framework offering browser-based assembly, configuration, deployment, fine-tuning, and local data incorporation.

reworkd/AgentGPT framework
Open GitHub

Auto-GPT is discussed as an open-source agent template/framework for decomposing objectives and executing tasks in a loop.

Significant-Gravitas/Auto-GPT framework
Open GitHub

SmolModels / smol-ai developer is listed as a code-domain LLM-based agent system with self-feedback and tool use.

smol-ai/developer system
Open GitHub

WorkGPT is discussed as an open-source GPT agent framework for invoking APIs.

team-openpm/workgpt framework
Open GitHub

SuperAGI is listed as a universal open-source autonomous AI agent framework.

TransformerOptimus/SuperAGI framework
Open GitHub

XLang is discussed as an open-source framework for building and evaluating language model agents through executable language grounding.

xlang-ai/xlang framework
Open GitHub

BabyAGI is listed as an LLM-based agent that creates tasks from objectives and stores or retrieves task results.

yoheinakajima/babyagi system
Open GitHub

Repository for the CogMir framework's core data assets and high-level experimental prompt-design logic, including datasets corresponding to Appendix C and prompt templates corresponding to Section 4 and Appendix D.

XuanL17/CogMir
Open GitHub

BabyAGI is described as an OpenAI-powered task management system that uses vector databases such as Chroma or Weaviate to manage, prioritize, execute, store, and recall task-related information.

yoheinakajima/babyagi framework
Open GitHub

Code for defining text-to-text tasks and mixtures, preprocessing and evaluating datasets, training and fine-tuning T5 models, and reproducing the paper's experiments, with links to released checkpoints.

google-research/text-to-text-transfer-transformer
Open GitHub

Repository for ExpNote: Black-box Large Language Models are Better Task Solvers with Experience Notebook, including experiment code, datasets, scripts, and setup instructions.

forangel2014/ExpNote dataset
Open GitHub

Repository containing the 50-claim development dataset, the 25-claim post-finalization dataset, and dataset statistics used in the paper.

LaraHack/linkflows_claims_dataset
Open GitHub

Repository containing materials from Stages 1, 2, and 3 of the expert formalization study evaluating the super-pattern.

LaraHack/linkflows_formalization_study
Open GitHub

Repository containing the paper's released Self-Contrast code and data.

THUDM/Self-Contrast implementation
Open GitHub

Open-source implementation of the paper's GPT-2 generation, candidate ranking, and training-data memorization attack workflow.

ftramer/LM_Memorization
Open GitHub

Repository containing FactReview, RefCopilot, demos, source code, scripts, tests, and documentation for evidence-grounded reviews of ML papers.

DEFENSE-SEU/FactReview framework
Open GitHub

Repository containing pilot and main JSON datasets, prompt generation code, evaluation code, prompt templates, and unit-group definitions for the FAITH paper.

ZHANG-MENGAO/FAITH dataset
Open GitHub

Author-linked PyTorch codebase for the paper's transformer self-supervision experiments, including SynCo-style synthetic negatives, DeiT/Swin configurations, pretraining, and linear evaluation.

giakoumoglou/synco-v2
Open GitHub

PyTorch code, configurations, training and unmasking scripts, evaluation utilities, and pretrained checkpoints for ImageNet 256x256 and 512x512 MaskDiT models.

Anima-Lab/MaskDiT implementation
Open GitHub

Code implementing the Favi-Score introduced and evaluated in the paper.

vodezhaw/faviscore
Open GitHub

Code repository linked by the paper for the FedSPM method and experiments.

zijianwang0510/FedSPM
Open GitHub

Official repository for reproducing the paper's output-refinement and policy-refinement experiments.

aypan17/llm-feedback
Open GitHub

Repository explicitly published with the paper for the FEVER annotation interfaces; the current repository also contains dataset preparation, scoring, and baseline components.

awslabs/fever
Open GitHub

Repository explicitly published with the paper for the FEVER baseline pipeline implementing evidence retrieval and textual entailment.

sheffieldnlp/fever-baselines
Open GitHub

Repository associated with the Cross-Attentive Time-Series Trend Network described and evaluated in the paper.

kieranjwood/x-trend framework
Open GitHub

Official repository for the paper, containing code for synthetic hallucination generation, reward-model training and evaluation, sample hallucination data, and links to released data.

du-nlp-lab/FG-PRM
Open GitHub

Open-source repository associated with PIXIU and FinBen, containing financial LLM resources, evaluation datasets, benchmark materials, code, and links to related models and leaderboards.

The-FinAI/PIXIU dataset
Open GitHub

Repository implementing FinBERT and providing training, prediction, and model resources associated with the paper.

ProsusAI/finBERT
Open GitHub

Repository indicated by the authors for the FinBERT2 work, including the specialized encoder and related downstream variants or resources.

ValueSimplex/FinBERT2 system
Open GitHub

ProgramFC repository used as the implementation basis for the LLM-based composite fact-checking system in the experiment.

teacherpeterpan/ProgramFC system
Open GitHub

Repository for FAVOR with model files, inference code, environment specification, example data, and checkpoint setup instructions.

BriansIDP/AudioVisualLLM
Open GitHub

Official repository providing the paper's Fine-Grained RLHF implementation, QA-FEEDBACK data, reward-model training code, RLHF scripts, trained-model references, and annotation interfaces.

allenai/FineGrainedRLHF
Open GitHub

Repository containing preprocessing, model-definition, training, testing, and performance-analysis code for mapping CSU veterinary medical summaries to SNOMED-CT diagnosis codes.

adam-kiehl/DiagnosisCoding
Open GitHub

Official repository for offline reward-model training and preference-based language-model fine-tuning, including smaller fine-tuned models and a subset of collected human labels.

openai/lm-human-preferences
Open GitHub

Official repository implementing the paper's RL4VLM training pipeline, GymCards environment, modified LLaVA components, and PPO training for GymCards and ALFWorld.

RL4VLM/RL4VLM
Open GitHub

GitHub repository for a Chinese financial news sentiment classification dataset containing train and test CSV files and financial news sentiment labels.

wwwxmu/Dataset-of-financial-news-sentiment-classification dataset
Open GitHub

Repository path containing the AdvBench harmful-behaviors data from which the paper uses a 50-prompt subset across 32 categories.

llm-attacks/llm-attacks
Open GitHub

Repository containing code for FinEAS and the paper's BERT, BiLSTM, and FinBERT financial-news sentiment experiments, together with reported result tables.

lhf-labs/finance-news-analysis-bert benchmark
Open GitHub

Repository containing data preprocessing code, FineFT algorithm training/validation/testing scripts, VAE routing components, trading-environment implementation, baseline code, and analysis utilities.

qinmoelei/FineFT_code_space framework
Open GitHub

Google Research Big Vision repository containing the UViM codebase used for the paper's depth, panoptic-segmentation, and colorization evaluations.

google-research/big_vision
Open GitHub

Official Google Research directory containing the FSQ paper README and reference FSQ Colab/code artifact.

google-research/google-research
Open GitHub

Official MaskGIT JAX implementation used as the basis for the paper's MaskGIT FSQ-versus-VQ experiments.

google-research/maskgit
Open GitHub

Python source code repository for FINMEM, the LLM trading agent with layered memory and character design.

pipiku915/FinMem-LLM-StockTrading system
Open GitHub

Repository for the FinReport code and datasets released by the paper.

frinkleko/FinReport dataset
Open GitHub

Astock dataset used for stock and news data, train-validation-test splitting, OOD analysis, and backtesting.

JinanZou/Astock dataset
Open GitHub

Repository for the paper's dataset construction pipeline, data collection module, FinRpt framework modules, benchmark evaluation code, fine-tuning setup, reinforcement-learning setup, and website front-end code.

jinsong8/FinRpt dataset
Open GitHub

GitHub repository associated with the paper's benchmark for online financial RAG evaluation.

PhealenWang/financial_rag_benchmark dataset
Open GitHub

Repository released by the authors for FinTexTS framework code and pilot study implementation.

leejaehoon2016/FinTexTS dataset
Open GitHub

Open-source implementation repository for FinWorld, the end-to-end financial AI research and deployment platform introduced in the paper.

DVampire/FinWorld framework
Open GitHub

Contains FireAct prompts, task and tool definitions, trajectory-generation and evaluation code, fine-tuning scripts, example training data, and model references.

anchen1011/FireAct implementation
Open GitHub

Open-source implementation of FlashAttention referenced by the paper.

HazyResearch/flash-attention implementation
Open GitHub

Source code and experiment resources for the Fleet of Agents framework introduced and evaluated in the paper.

au-clan/FoA
Open GitHub

Repository containing the Flow multi-agent workflow automation implementation, including workflow management code, prompts, validators, notebooks, and generated examples.

tmllab/2025_ICLR_FLOW framework
Open GitHub

Code repository for the FlowAgent framework introduced by the paper.

Lightblues/FlowAgent framework
Open GitHub

Official repository containing FlowBench source data organization, turn-level and session-level evaluation code, scripts, prompts, and setup instructions.

Justherozen/FlowBench dataset
Open GitHub

GitHub Gist with a replenishment plan generated by the Flowr DC Replenishment Planning Agent, including outlet allocations, route assignments, vehicle assignments, route summary, consolidation opportunities, and human review queue.

https:/
Open GitHub

GitHub Gist with a purchase-order report generated by the Flowr Procurement and Ordering Agent, including order quantities, supplier justifications, delivery estimates, consolidated supplier orders, and human review flags.

https:/
Open GitHub

Repository listed by the paper for code associated with dataset analysis, object-detection benchmarking, and cost-aware sequence active-learning experiments.

olivesgatech/FOCAL dataset
Open GitHub

Repository made available by the authors for the proposed models used to forecast extreme Bitcoin volatility movements.

dorienh/bitcoin_synthesizer implementation
Open GitHub

Repository for the FiGASR package and related sentiment-indicator resources.

lucabarbaglia/FiGASR package
Open GitHub

Repository containing Python code for fine-grained aspect-based sentiment analysis in the economic news setting.

sergioconsoli/SentiBigNomics implementation
Open GitHub

Public multivariate time-series dataset repository containing the exchange-rate data used as real-data-E.

laiguokun/multivariate-time-series-data dataset
Open GitHub

Lag-Llama pretrained model repository used as the probabilistic zero-shot forecaster in the experiments.

time-series-foundation-models/Lag-Llama implementation
Open GitHub

Official implementation of the Diffusion Model-Based Predictor baseline for robust offline RL against state-observation perturbations.

zhyang2226/DMBP implementation
Open GitHub

Repository containing the implementation and results for LSTM and GRU time-series forecasting experiments discussed by the paper.

Alebuenoaz/LSTM-and-GRU-Time-Series-Forecasting implementation
Open GitHub

Repository containing FoT code, task and benchmark scripts, datasets, model and method modules, and evaluation instructions for Game of 24, GSM8K, MATH-500, and AIME.

iamhankai/Forest-of-Thought framework
Open GitHub

Repository containing more detailed prompts and implementation code for the LAWN antenna self-evolution framework discussed in the paper.

ChangyuanZhao/LAE_evolving framework
Open GitHub

Repository stated by the authors to contain the complete source code employed in the research.

Suhasnadh/Power-Consuption implementation
Open GitHub

Official code for the paper, including the local Qwen3-4B/BGE-M3 RAG pipeline, controlled diagnostic experiment, BEIR evaluation, cross-encoder comparison, and gating analyses.

Silk-Road/causal-rag-rerank
Open GitHub

Repository for the paper 'From Coarse to Fine: Self-Adaptive Hierarchical Planning for LLM Agents', containing implementation-related folders such as agents, configs, envs, plan generation, prompts, tasks, utilities, and experiment scripts.

import-myself/AHP framework
Open GitHub

Repository for the paper From Emergence to Control: Probing and Modulating Self-Reflection in Language Models, describing probing vectors, model insertion, and reflection analysis for controlling self-reflection in LLMs.

xzAscC/ProbingReflection implementation
Open GitHub

Repository for ALARM, including the front-end GUI, back-end algorithms, and the anomaly/feature-importance simulation code described by the authors.

xyvivian/ALARM
Open GitHub

A maintained paper list and resource collection about LLM-as-a-judge.

llm-as-a-judge/Awesome-LLM-as-a-judge
Open GitHub

Repository for the paper's semi-automated ontology and knowledge-graph construction pipeline, including prompts, code, data, generated artifacts, results, and evaluation materials.

fusion-jena/automatic-KG-creation-with-LLM
Open GitHub

Repository containing the reproducibility-related publication data and analysis code from the authors' prior biodiversity deep-learning study, which supplied the source dataset for the current pipeline.

fusion-jena/Reproduce-DLmethods-Biodiv
Open GitHub

Repository accompanying the review, containing the detailed thematic scoring by co-authors, supplementary review material, and diagrams.

mak-raiaan/LLMAgentsReview
Open GitHub

Repository titled for the paper 'From Large AI Models to Agentic AI: A Tutorial on Future Intelligent Communications' and identified by GitHub as the code repository for the paper.

jiangfeibo/ComAgent framework
Open GitHub

BeeAI is described as the experimental platform central to IBM's ACP, supporting local-first orchestration, agent discovery, REST endpoints, SDKs, telemetry, and multi-agent execution.

i-am-bee/beeai-framework framework
Open GitHub

The MCP servers repository is cited as an ecosystem of reference and integration servers for file management, databases, Google Drive, Git, GitHub, GitLab, Slack, Google Maps, image generators, and search APIs.

modelcontextprotocol/servers implementation
Open GitHub

OpenAI Swarm is reviewed as a lightweight, stateless abstraction for multi-agent systems with agent definitions, dynamic handoffs, context management, direct function calling, streaming, and backend flexibility.

openai/swarm framework
Open GitHub

Open-source Microsoft GraphRAG repository implementing the graph-based indexing and retrieval/summarization approach described in the paper.

microsoft/graphrag implementation
Open GitHub

Repository reported by the paper as containing code and data for the proposed news-aware LLM time series forecasting framework.

ameliawong1996/From_News_to_Forecast dataset
Open GitHub

Open-source OpenCode repository corresponding to one of the agentic tools evaluated by the paper.

anomalyco/opencode
Open GitHub

Open-source Kilo Code repository corresponding to the agentic coding tool selected for the paper's second prototype.

Kilo-Org/kilocode
Open GitHub

Repository containing code and data for the HippoRAG family, including the framework introduced and evaluated in this paper.

OSU-NLP-Group/HippoRAG framework
Open GitHub

Repository path for a Workflow Composer that transforms natural-language research questions into executable HyperFlow workflows for the 1000 Genomes Project, including source code, tests, Skills/resources, CLI commands, and evaluation-related materials.

hyperflow-wms/1000genome-workflow dataset
Open GitHub

A companion GitHub repository associated with the survey on agentic workflow optimization.

IBM/awesome-agentic-workflow-optimization
Open GitHub

Repository for the paper containing raw experimental data, reasoning processes and outputs, keyword-statistic results, LLM-as-a-judge scores, and score-analysis spreadsheets.

ChangWenhan/FromThinking2Output dataset
Open GitHub

Official companion repository for the tutorial, containing the paper materials and a resource browser organized around its world-model and world-action-model taxonomy.

clearlab-sustech/WorldModelSurvey
Open GitHub

Repository associated with the paper<ef><bf><bd>s recursors/refiners work, containing generated Prolog programs, models, execution traces, and related materials for the implemented system.

ptarau/recursors framework
Open GitHub

AWorld-RL is the linked project repository that houses FunReason-MT materials and describes the introduced multi-turn function-calling data-synthesis framework.

inclusionAI/AWorld-RL framework
Open GitHub

The FunReason v1 paper explicitly directs readers to this GitHub repository for code and dataset; the URL currently redirects to the renamed BalanceSFT repository.

BingguangHao/FunReason dataset
Open GitHub

Repository identified by the paper as containing the pre-processed dataset, raw news article data, and implementation code for M2VN.

Yoontae6719/M2VN-Multi-Modal-Learning-Network-for-Volatility-Forecasting dataset
Open GitHub

Repository containing GAAMA's typed memory graph, semantic and PPR retrieval components, prompts, storage adapters, and LoCoMo evaluation scripts.

swarna-kpaul/gaama framework
Open GitHub

Repository containing retrieval, gain-signal synthesis, selector training, and evaluation code for GainRAG.

liunian-Jay/GainRAG
Open GitHub

Repository for the Game-theoretic LLM paper, including source code, setup instructions, complete-information game experiments, workflow experiments, and Deal-or-No-Deal experiments.

Wenyueh/game_theory benchmark
Open GitHub

Official GameplayQA benchmark codebase containing data-processing, model-evaluation, judging, ablation, evaluation, and plotting workflows.

HATS-ICT/GameplayQA benchmark
Open GitHub

Web-based tool developed for synchronized single- and multi-video timeline annotation and verification used in constructing GameplayQA.

wangyz1999/sync-video-label
Open GitHub

Gapetron is the modular large-scale pretraining toolkit used to train the Gaperon model suite and released as a primary implementation artifact of the paper.

NathanGodey/gapetron
Open GitHub

Repository reported by the author as containing the code and dataset used in the paper.

bvsdinda/TPE-LSTM dataset
Open GitHub

Official implementation and evaluation code for GEAR, including generation benchmarks and a CUDA-supported GEAR-KIVI implementation.

opengear-project/GEAR
Open GitHub

Repository containing the code used to construct and evaluate nearest-neighbor language models, including datastore building and FAISS-backed retrieval.

urvashik/knnlm implementation
Open GitHub

ytopt is a machine-learning-based autotuning and hyperparameter-optimization framework used in Section IV-B to search coefficient spaces for the reviewed scaling laws.

ytopt-team/ytopt package
Open GitHub

The Lila benchmark repository supplies the program-form mathematical reasoning data used in the paper's mathematical program-synthesis experiments.

allenai/Lila dataset
Open GitHub

Official AASbyLLM research repository containing the prototype implementation, simplified source code, AAS samples, prompts, raw evaluation results, evaluation spreadsheet, figures, and demos associated with the paper.

YuchenXia/AASbyLLM
Open GitHub

Repository released by the authors for the GAR method, including code and retrieval results.

morningmoni/GAR
Open GitHub

Public repository for the paper's generative-agent architecture and Smallville simulation code.

joonspk-research/generative_agents
Open GitHub

Repository linked directly by the paper; it provides the satellite-domain dataset/knowledge material and a simple notebook example for validating the generative AI agent.

RickyZang/GAI-agent-satellite dataset
Open GitHub

Official repository for the ToxiGen synthetic toxic/benign text dataset evaluated as an alternative and complementary augmentation source.

microsoft/ToxiGen dataset
Open GitHub

PyTorch repository containing generative sparse-index-tracking code, comparison code, a GECCO 2023 directory, backtesting material, and instructions for obtaining the associated dataset.

kayuksel/generative-opt implementation
Open GitHub

Code for training WSGAN with a StyleGAN2-ADA base architecture on weakly supervised CIFAR10 and LSUN datasets.

benbo/stylewsgan
Open GitHub

Code for jointly training the DCGAN-based WSGAN and label model, together with weak-label files and instructions for the principal datasets.

benbo/WSGAN-paper
Open GitHub

Public repository from which the authors select 96 model-specific jailbreak prompts used as the jailbreak query group.

elder-plinius/L1B3RT4S
Open GitHub

Official code repository for GRT, including task configurations, VLM-based keypoint preparation, population-based deformation search, simulator evaluators, and analysis scripts.

RCHI-Lab/GRT
Open GitHub

Repository explicitly linked by the paper for GestureGPT prompts and corresponding data samples used to guide the three-agent architecture.

studyzx/GestureGPT_ISS
Open GitHub

Official source repository for the Gibson Environment, including installation, assets integration, rendering, agent examples, and training demos.

StanfordVL/GibsonEnv dataset
Open GitHub

Official repository containing Gorilla inference resources, APIBench data, evaluation scripts, model outputs, and materials for reproducing the paper's results.

ShishirPatil/gorilla
Open GitHub

OpenAI Evals, the framework the report states it is open-sourcing for creating and running model benchmarks and inspecting performance sample by sample.

openai/evals
Open GitHub

Repository for the SeeAct generalist web agent, including code, data links, and evaluation tools associated with the paper.

OSU-NLP-Group/SeeAct
Open GitHub

Repository containing code and instructions for preparing nuScenes-derived prompts, fine-tuning GPT-3.5, generating motion-planning outputs, and evaluating the GPT-Driver approach.

PointsCoder/GPT-Driver implementation
Open GitHub

Public prompts and code for the GPT-based automated review generation workflow used in the study.

zrobertson466920/GPT_Auto_Review implementation
Open GitHub

Author-owned repository containing ROG experiment code, FedAvg attack scripts, compression configurations, model components, and minimal example data.

KAI-YUE/rog
Open GitHub

Repository linked by the paper as the implementation of Grammar Search for multi-agent systems.

mayanks43/grammar_search framework
Open GitHub

Repository providing the authors' implementation of a Graph Attention Network layer and an execution example.

PetarV-/GAT implementation
Open GitHub

Code and sample artifacts for concept-specific KG generation, UMLS subgraph sampling, node and edge clustering, personalized KG composition, BAT model training, prediction, and baseline execution.

pat-jj/GraphCare implementation
Open GitHub

Source code for GraphGPT and the Graph Eulerian Transformer workflow.

alibaba/graph-gpt framework
Open GitHub

RAGChecker is an open-source visual analytics tool that compares LLM outputs against source documents, supports GraphEval+ and SICI-style detection, and presents claim reliability through an interactive quadrant-based visualization.

tanmayagrawal21/RAGChecker framework
Open GitHub

Open R1 provides the GRPO/SFT training framework that the study extends and is also the source of the no-reflection comparison used in the reflection-reward ablation discussion.

huggingface/open-r1
Open GitHub

Repository supplied by the author for the paper's GRPO-with-reflection implementation, training workflow, evaluation code, and reported experimental results.

Red-Scarff/GRPO_reflection implementation
Open GitHub

Repository accompanying the paper with GSM-Symbolic templates and generated benchmark data, including GSM-Symbolic, GSM-Symbolic-P1, and GSM-Symbolic-P2 resources.

apple/ml-gsm-symbolic dataset
Open GitHub

Repository for the GTA benchmark, dataset, code, and evaluation materials for general tool agents.

open-compass/GTA dataset
Open GitHub

Repository for the GUARD-SLM method, including activation extraction and analysis code, activation classification, inference and judging scripts, datasets, and experimental results.

solidlabnetwork/GUARD-SLM
Open GitHub

Official GuardAgent code repository containing the guard agent implementation, prompts, callable tools, and runner scripts.

guardagent/code implementation
Open GitHub

Official repository for the EICU-AC and Mind2Web-SC benchmark resources introduced by the paper.

guardagent/dataset dataset
Open GitHub

Official repository for H3M-SSMoEs containing Python code, model components, training and backtesting scripts, and links to datasets and model weights.

PeilinTime/H3M-SSMoEs dataset
Open GitHub

Repository explicitly linked by the paper for the HaGRID dataset, downsampled/trial versions, pretrained models, and dynamic gesture recognition demo.

hukenovs/hagrid dataset
Open GitHub

Official repository for the HalluLens benchmark, including materials for the extrinsic hallucination evaluation tasks and associated evaluation workflow.

facebookresearch/HalluLens
Open GitHub

The FacTool repository contains a dedicated Halu-J directory and the resources released for the critique-based hallucination judge and ME-FEVER.

GAIR-NLP/factool
Open GitHub

Official open-source implementation and data repository for HarmBench, including benchmark data, red-teaming baselines, evaluation classifiers, pipeline scripts, and adversarial-training code.

centerforaisafety/HarmBench benchmark
Open GitHub

Repository stated as the release location for HDFlow code and data.

wenlinyao/HDFlow dataset
Open GitHub

Documentation and examples for obtaining embeddings from the HeAR health-acoustic foundation model through its research API.

Google-Health/google-health
Open GitHub

Repository directory containing CT Foundation API documentation, demos, and resources for generating embeddings from CT volumes.

Google-Health/imaging-research
Open GitHub

Repository associated with the paper's released LOFin benchmark and HiREC implementation.

deep-over/LOFin-bench-HiREC dataset
Open GitHub

Official repository linked by the paper for HippoRAG code and data.

OSU-NLP-Group/HippoRAG
Open GitHub

Hosts the paper project page, benchmark overview, examples, statistics, comparison graphics, and model leaderboard for HLV-1K.

Vincent-ZHQ/HLV-1K dataset
Open GitHub

Open-source implementation of the HAL evaluation harness introduced and evaluated in the paper.

princeton-pli/hal-harness
Open GitHub

Open-source HomeRobot codebase containing shared agents and interfaces, Habitat-based simulation support, hardware control for the Stretch robot, and OVMM baseline implementations.

facebookresearch/home-robot
Open GitHub

Repository containing code to reproduce the HCNN experiments reported in the paper.

FinancialComputingUCL/HomologicalCNN benchmark
Open GitHub

PyTorch implementation of the Hopfield, HopfieldPooling, and related continuous modern Hopfield-layer functionality introduced and evaluated in the paper.

ml-jku/hopfield-layers
Open GitHub

Repository containing the baseline model code and pipeline for HotpotQA data download, preprocessing, training, prediction, and evaluation.

hotpotqa/hotpot
Open GitHub

Repository containing annotated transcripts of each participant made available for future studies.

safety-research/how-ai-impacts-skill-formation dataset
Open GitHub

TheAgentCompany repository contains sandboxed work environments, task directories, evaluators, task instructions, and supporting files for many task instances listed in the paper's appendix task table.

TheAgentCompany/TheAgentCompany dataset
Open GitHub

Repository for coding-agent token-consumption analysis, including dataset-building scripts, multi-model analysis scripts, phase-level token decomposition, and self-prediction correlation computation.

LongjuBai/agent_token_consumption_analysis implementation
Open GitHub

Python codebase for running the ensemble generalization-gap experiments, computing dataset complexity metrics, and generating per-dataset outputs and figures.

zubair0831/ensemble-generalization-gap benchmark
Open GitHub

Open-Instruct repository released with the paper to support reproduction and further research on open instruction tuning, including T"ULU-related artifacts.

allenai/open-instruct
Open GitHub

GitHub directory linked by the paper for the deterministic prediction task source code, including Pauli string multiplication, divide-and-conquer, letter replacement, and addition-related files.

EdenCodeInc/PyCliffordMCP benchmark
Open GitHub

Repository containing the dataset and experimental code for the study of memory addition, deletion, and experience-following behaviour in LLM agents.

yuplin2333/agent_memory_manage dataset
Open GitHub

Repository containing the prototype implementation for customizing and probing leaderboards with the paper's evaluation approach.

swarooprm/Leaderboard-Customization
Open GitHub

The repository provides code folders for HAG-XAI object detection, HAG-XAI image classification, FullGradCAM for Yolo-v5s, and FullGradCAM for Faster-RCNN, along with links to experimental materials, human attention data, and pretrained model files.

GitVirTer/HAG-XAI framework
Open GitHub

Repository linked directly from the paper; it contains the implementation notebook, Appendix.pdf, Focus data, and transformed MLC data.

ruijiang81/hai-blbf implementation
Open GitHub

Repository containing data and analysis code for the human-alignment AI-assisted decision-making study.

Networks-Learning/human-alignment-study dataset
Open GitHub

Contains the Auto_Driving_Highway code, prompt materials, results, and instructions for training an RL agent with an LLM in the reward loop.

JingYue2000/In-context_Learning_for_Automated_Driving framework
Open GitHub

Azure Verified Modules is used by the paper as a cloud-native curated module ecosystem that demonstrates governance, standards, testing, versioning, and consistent interfaces.

Azure/Azure-Verified-Modules framework
Open GitHub

Repository indicated by the paper for prompts, code, anonymised data, and the full questionnaire related to the AI Narrative Test.

Mosh0110/AI-narrative-test benchmark
Open GitHub

Repository for the S&P 500 membership forecasting project, including EDA.ipynb, Modeling_Process.ipynb, model-interpretability figures, and Python dependencies.

VidhiAgrawal/sp500_finance implementation
Open GitHub

Source code for the HybRank hybrid and collaborative passage-reranking model.

zmzhang2000/HybRank implementation
Open GitHub

Repository linked by the paper as the code for the HyGRL framework.

wjywjy123/HyGRL implementation
Open GitHub

Repository containing the HypeLoRA MLP and Transformer hyper-networks, dynamic LoRA layers, RoBERTa builders, GLUE data loading, calibration metrics, experiment configurations, and experiment-running utilities.

btrojan-official/HypeLoRA implementation
Open GitHub

Official implementation repository for HyperAgent, including the multi-agent software engineering framework and benchmark reproduction scripts.

FSoft-AI4Code/HyperAgent framework
Open GitHub

GitHub repository for the ICICLE method introduced and evaluated in the paper.

gmum/ICICLE framework
Open GitHub

Repository containing ToolEmu code, emulators, evaluators, curated toolkit and test-case assets, scripts, and notebooks.

ryoungj/ToolEmu benchmark
Open GitHub

Repository for the IDQL method introduced in the paper, including offline-training and fine-tuning scripts and the diffusion/IQL learner implementation.

philippe-eecs/IDQL implementation
Open GitHub

Author-released implementation of PromptInject, the modular framework used to assemble prompts and evaluate LLM robustness to adversarial prompt attacks.

agencyenterprise/PromptInject
Open GitHub

GitHub repository released by the authors for IL-PCSR and developed models.

Exploration-Lab/IL-PCSR dataset
Open GitHub

Repository containing data preparation, BERT-based and T5 training and evaluation scripts, late-fusion evaluation, and experiment configurations for the paper's CBM models.

salanueva/CBM
Open GitHub

Repository containing the ImageRAG codebase, configurations, scripts, evaluation workflow, and pointers to released data, caches, and checkpoints.

om-ai-lab/ImageRAG implementation
Open GitHub

Complete codebase and setup instructions for the sentiment-driven stock prediction experiments.

Walids35/capstone-stock-prediction implementation
Open GitHub

Repository maintained by the authors as an official page and living resource list for the survey on implicit reasoning in LLMs.

digailab/awesome-llm-implicit-reasoning
Open GitHub

Official HarmBench repository used to generate adversarial safety-evaluation test cases for the downstream RLHF experiment.

centerforaisafety/HarmBench
Open GitHub

Open-source repository cited as the source for the Reuters & Bloomberg news-title stock-prediction dataset used to build headline vectors.

WenchenLi/news-title-stock-prediction-pytorch dataset
Open GitHub

Repository titled Anote-Text-Classification containing dataset folders, trial runs, requirements, README explanations, and model performance comparisons for GPT-3.5 Turbo, SetFit, and BERT across the evaluated datasets.

iiWhiteii/Anote-Text-Classification implementation
Open GitHub

Repository for the CARE native retrieval-augmented reasoning framework, including scripts, evaluation resources, documentation, and training examples.

FoundationAgents/CARE framework
Open GitHub

Official FactCC implementation and model used to score whether generated summary claims are factually consistent with source documents.

salesforce/factCC implementation
Open GitHub

Preliminary implementation of the paper's multiagent debate experiments, with code for arithmetic, GSM8K, biography, and MMLU tasks.

composable-models/llm_multiagent_debate
Open GitHub

Repository for Fusion-in-Decoder, providing the Natural Questions version augmented with DPR-retrieved passages that RETRO uses for its question-answering experiment and representing a principal comparison system.

facebookresearch/FiD
Open GitHub

NexusRaven-V2 repository linked by the paper for the Nexus function-calling evaluation benchmark.

nexusflowai/NexusRaven-V2 benchmark
Open GitHub

Repository containing notebooks for SFT and LoRA-based DPO training, generation and reward-model evaluation of the four OPT-350M variants, evaluation data and JSON outputs, plots, and examples of noisy preference pairs.

PiyushWithPant/Improving-LLM-Safety-and-Helpfulness-using-SFT-and-DPO dataset
Open GitHub

Code repository for the paper's QE-based machine-translation feedback-training experiments and RAFT+ method.

zwhe99/FeedbackMT implementation
Open GitHub

Code repository for the Synchronously Self-Reviewing fine-tuning method introduced and evaluated in the paper.

liangyupu/SSR
Open GitHub

Repository for the Unsupervised Passage Re-ranking method introduced in the paper, including the codebase and paper-linked data/checkpoints.

DevSinghSachan/unsupervised-passage-reranking
Open GitHub

Code repository for the MAP paper, including implementations and evaluation scripts for Tower of Hanoi, CogEval graph tasks, PlanBench, StrategyQA, and transfer experiments.

MAPLLM/MAPICLR2025sub system
Open GitHub

Microchain is the agent-based library used to inject function declarations, descriptions, and examples into prompts, call functions during the reasoning loop, and record assistant/user chat-completion chains.

galatolofederico/microchain framework
Open GitHub

S-LoRA is the scalable multi-adapter serving baseline whose unified pool does not retain history KV caches in the evaluated configuration.

S-LoRA/S-LoRA
Open GitHub

vLLM is the serving engine extended by FastLibra and a principal baseline with static LoRA/KV cache partitioning and LRU management.

vllm-project/vllm
Open GitHub

Repository containing code and standardized data for direct experiments, indirect matched-content experiments, the AllSides news-choice case study, and the Amazon seller-choice case study; it also documents the inference environment and points to hosted aggregated outputs.

aflah02/LLM-Latent-Source-Preferences dataset
Open GitHub

Public repository for InlineCoder, the framework introduced and evaluated in the paper.

ythere-y/InlineCoder dataset
Open GitHub

Repository containing the MMT-VQA Fairseq-based implementation, training and evaluation scripts, dataset preparation instructions, and access information for Multi30K-VQA.

libeineu/MMT-VQA
Open GitHub

Official repository listed by the paper for the InfiR model family and associated model resources.

Reallm-Labs/InfiR implementation
Open GitHub

Repository linked by the paper for InfoMosaic-Bench/InfoMosaic-Flow resources.

DorothyDUUU/Info-Mosaic dataset
Open GitHub

Published ETT dataset collected by the authors and used as a core benchmark dataset in the paper.

zhouhaoyi/ETDataset dataset
Open GitHub

Source code for the Informer model and experiments.

zhouhaoyi/Informer2020 framework
Open GitHub

Aider source repository analyzed for user-driven loop design, repo-map retrieval, edit formats, and summarization behavior.

Aider-AI/aider implementation
Open GitHub

OpenCode source repository analyzed for tool interface design, event bus, SQLite session persistence, dynamic tools, and role-based sub-agents.

anomalyco/opencode implementation
Open GitHub

Moatless Tools source repository analyzed for MCTS orchestration, tree-structured state, action classes, semantic search, in-memory shadow mode, and actor-critic routing.

aorwall/moatless-tools implementation
Open GitHub

AutoCodeRover source repository analyzed for phased scaffold control, search-only tools, AST-aware retrieval, SBFL, and Docker-based execution.

AutoCodeRoverSG/auto-code-rover implementation
Open GitHub

Cline source repository analyzed for recursive control flow, IDE coupling, shadow git checkpoints, delegation, and LLM-initiated compaction.

cline/cline implementation
Open GitHub

Prometheus source repository analyzed for LangGraph-based phased control, graph-scoped state, knowledge graph retrieval, per-node tool scoping, and multi-tier persistence.

EuniAI/Prometheus implementation
Open GitHub

Gemini CLI source repository analyzed as one of the 13 coding agent scaffolds.

google-gemini/gemini-cli implementation
Open GitHub

Codex CLI source repository analyzed for event-driven ReAct control, dynamic tool rebuilding, sandboxing, Guardian safety routing, memory extraction, and sub-agent delegation.

openai/codex implementation
Open GitHub

Agentless source repository analyzed for fixed-pipeline architecture, hierarchical localization, sampling, and JSONL pipeline state.

OpenAutoCoder/Agentless implementation
Open GitHub

OpenHands source repository analyzed for event-sourced architecture, tool interfaces, Docker execution, and delegation mechanisms.

OpenHands/OpenHands implementation
Open GitHub

mini-swe-agent source repository analyzed as a deliberately minimal baseline scaffold with a single bash tool and simple ReAct loop.

SWE-agent/mini-swe-agent implementation
Open GitHub

SWE-agent source repository analyzed for ReAct control loop, tool bundles, retry behavior, Docker isolation, and compaction processors.

SWE-agent/SWE-agent implementation
Open GitHub

DARS-Agent source repository analyzed for depth-first tree search, SWE-agent-derived tools, greedy LLM critic selection, and Docker reset/replay state recovery.

vaibhavagg303/DARS-Agent implementation
Open GitHub

A topic-agnostic, provenance-first pipeline for legitimate AI-assisted drafting of a Related Work section using human-authored structured notes, strict no-invention constraints, interaction logs, generated taxonomy, draft output, audit table, provenance card, and related artifacts.

RutaBinkyte/AI-RO framework
Open GitHub

Repository associated with the paper's survey of instruction-tuning research.

xiaoya-li/Instruction-Tuning-Survey
Open GitHub

Python repository associated with hybrid physics-based and data-driven building energy modeling, with folders for evaluation, feature generation, forecasting, models, preprocessing, and utilities.

Leo-VK/hybrid_bem dataset
Open GitHub

Official DVLR repository containing training, MathVerse evaluation, example, and environment files for the proposed decoupled visual interpretation and linguistic reasoning framework.

guozix/DVLR implementation
Open GitHub

Repository containing code notebooks, data files, tickers, metrics/metadata, SQL queries, and the list of companies associated with the paper's algorithmic stock-market trading system.

JuanCarlosKing/StockmarketAlgoritmicTrading dataset
Open GitHub

Repository for the MiniGrid gridworld environment used for the paper's four-room and six-room navigation experiments.

maximecb/gym-minigrid benchmark
Open GitHub

Repository for the FastAPI control server, callback-augmented interactive trainer, React/TypeScript dashboard, examples, and LLM-based tuning demonstration.

yuntian-group/interactive-training framework
Open GitHub

The paper states that the ICM protocol is open source under the MIT license and that referenced workspaces are available or buildable through this repository.

RinDig/Interpretable-Context-Methodology-ICM- framework
Open GitHub

Repository identified by the paper as containing the code for the GDELT headline extraction, FinBERT sentiment scoring, feature engineering, modeling, and backtesting workflow.

yukepenn/macro-news-sentiment-trading implementation
Open GitHub

Public code repository for the paper's context-faithfulness experiments.

liyp0095/ContextFaithful benchmark
Open GitHub

Repository for InvestLM, the LLaMA-based financial-domain instruction-tuned model released to the research community under the same licensing terms as LLaMA.

AbaciNLP/InvestLM implementation
Open GitHub

Python project for forecasting variance-covariance matrices in a Markowitz framework across cryptocurrency and traditional asset markets; the repository README states that it produced the paper's results.

Maciej-13/vcov_forecast framework
Open GitHub

Repository for IPO Finance Agent containing benchmark code, data directories, rubric artifacts, model-result files, and domain/workflow analysis outputs.

benstaf/ipoagent benchmark
Open GitHub

Repository containing the IQ-Learn implementation and usage materials for the paper's algorithm.

Div99/IQ-Learn
Open GitHub

Repository for the IRPAPERS benchmark introduced by the paper.

weaviate/IRPAPERS dataset
Open GitHub

Experimental code used for the retrieval and question-answering evaluations reported in the paper.

weaviate/query-agent-benchmarking
Open GitHub

Official Python repository accompanying the paper, organized into pretraining and fine-tuning code.

avinashsai/MML
Open GitHub

Open-source baseline compliance assessment agent for the CISO persona.

IBM/itbench-ciso-caa-agent
Open GitHub

Repository containing the sample ITBench scenarios released for community familiarization and development.

IBM/itbench-sample-scenarios
Open GitHub

Open-source baseline SRE agent introduced and evaluated with ITBench.

IBM/itbench-sre-agent
Open GitHub

Code and data repository for reproducing the paper's evaluator training, SFT+ILR, SFT+DPO, naive-ILR, and supervision-quality experiments.

helloelwin/iterative-label-refinement dataset
Open GitHub

Repository for JailbreakEval, an integrated toolkit containing multiple mainstream safety evaluators and supporting voting-based judgments of jailbreak success.

ThuCCSLab/JailbreakEval
Open GitHub

Repository storing submitted jailbreak prompts, responses, classifications, and metadata for attacks and defenses tracked by JailbreakBench.

JailbreakBench/artifacts
Open GitHub

Official JailbreakBench codebase implementing dataset access, model querying, jailbreak and refusal judges, red-teaming evaluation, defenses, logging, and submission workflows.

JailbreakBench/jailbreakbench benchmark
Open GitHub

Repository containing Stage 1 JEPA+DAAM training, Stage 2 decoder training, FSQ and mixed-radix packing, a HiFi-GAN decoder with optional DAAM gating, and DeepSpeed integration.

gioannides/Density-Adaptive-JEPA system
Open GitHub

Repository associated with the Jr. AI Scientist system; the paper points readers there for issues, comments, questions, and planned codebase release.

Agent4Science-UTokyo/Jr.AI-Scientist system
Open GitHub

FastChat's LLM-judge artifact containing the MT-bench evaluation machinery, judge prompts, and released benchmark/preference resources associated with the paper.

lm-sys/FastChat benchmark
Open GitHub

Official K-COMP codebase with retrieval, data-processing, training, inference scripts, environment configuration, short medical descriptions, and links to model checkpoints and datasets.

jeonghun3572/K-COMP framework
Open GitHub

Facebook/FAIR Memory Networks repository containing a KVmemnn directory with the Key-Value Memory Network model and scripts for building, training, evaluating, and interactively inspecting WikiMovies memory data.

facebook/MemNN implementation
Open GitHub

Repository path containing code or supplementary material for the shell-based keyword-search agent used in the paper.

amazon-science/aws-research-science system
Open GitHub

Repository for KG-Reasoner containing KG retrieval, GNN ranking, backtracking, RL training, reward server, evaluation scripts, and training/test data directories.

Wangshuaiia/KG-Reasoner
Open GitHub

Official implementation of the KnowAgent framework, including HotpotQA and ALFWorld path-generation scripts, trajectory filtering and merging, and LoRA-based knowledgeable self-learning.

zjunlp/KnowAgent
Open GitHub

Repository containing the question-generation compression workflow, paper-card generation, syntactic multihop experiments, datasets, evaluation scripts, and reported result files.

anvix9/llama2-chat system
Open GitHub

Official repository released with the paper, containing extensible implementations of KTO, DPO, offline PPO, ORPO, and other human-aware loss functions.

ContextualAI/HALOs implementation
Open GitHub

Repository for the technical report, providing code for high-frequency trading prediction under label imbalance, including MLP, LSTM, BERT, and Mamba backbones plus class weighting and data balancing options. The original dataset is not provided because of copyright restrictions.

RS2002/Label-Unbalance-in-High-Frequency-Trading framework
Open GitHub

Demonstration repository for the paper, including environment setup, dataset preparation, unsupervised detector execution, TS2Vec representation training, and feature preprocessing.

imTurkey/Label-Efficient-Interactive-Time-Series-Anomaly-Detection implementation
Open GitHub

Repository for the KPI anomaly-detection dataset used as one of LEIAD's three benchmark datasets.

NetManAIOps/KPI-Anomaly-Detection benchmark
Open GitHub

Official LaMMA-P repository containing code, PDDL resources, scripts, Fast Downward submodule setup, and MAT-THOR test dataset files for running the method in AI2-THOR.

tasl-lab/LaMMA-P dataset
Open GitHub

Open-source GPTSwarm implementation for building graph-based LLM agents and composite swarms, with modules for graph execution, LLM backends, memory, environments, and agent optimization.

metauto-ai/gptswarm
Open GitHub

Repository linked by the paper for code and data supporting Uncertainty Quantification with Attention Chain.

Yinghao-Li/UQAC
Open GitHub

Official repository accompanying the GPT-3 paper, containing synthetic arithmetic and word-scrambling datasets, dataset statistics, 175B samples, a model card, and benchmark-overlap examples.

openai/gpt-3
Open GitHub

Repository for PRONTOQA containing data-generation and analysis code plus released model outputs; the paper states that these artifacts reproduce its analysis and figures.

asaparov/prontoqa
Open GitHub

The repository contains an mgsm directory with TSV files for the ten translated languages plus English and manually translated few-shot exemplars.

google-research/url-nlp dataset
Open GitHub

Official demonstration code implementing the paper's language-model planning and admissible-action grounding workflow.

huangwl18/language-planner
Open GitHub

PIXIU provides FinMA models, FLARE financial evaluation benchmark tasks, and instruction data used as a baseline or source in the paper.

chancefocus/PIXIU benchmark
Open GitHub

Curated collection of papers associated with the survey's coverage of LLM-agent research.

luo-junyu/Awesome-Agent-Papers
Open GitHub

Open-source CAMEL framework for autonomous cooperation among communicative agents using inception prompting and role-play.

camel-ai/camel framework
Open GitHub

Open-source multi-agent collaborative framework associated with MetaGPT, discussed as a representative framework that embeds human workflow processes and SOPs into language-agent collaboration.

geekan/MetaGPT framework
Open GitHub

Open-source AutoGen framework for creating LLM applications using customizable agents that can be programmed through natural language and code.

microsoft/autogen framework
Open GitHub

Author-maintained repository for tracking LLM-based multi-agent papers and organizing them into streams such as frameworks, orchestration and efficiency, problem solving, world simulation, datasets, and benchmarks.

taichengguo/LLM_MultiAgents_Survey_Papers
Open GitHub

A GitHub repository made public by the authors to curate papers related to LLM safety.

tjunlp-lab/Awesome-LLM-Safety-Papers
Open GitHub

Repository containing the paper list associated with the survey on LLM-based agents for software engineering.

FudanSELab/Agent4SE-Paper-List dataset
Open GitHub

Official repository containing the paper's human-agent collaboration datasets, training/testing code for HotpotQA, StrategyQA, and InterCode, processed training data, and original evaluation outputs.

XueyangFeng/ReHAC dataset
Open GitHub

Repository containing the mathematical problem dataset, evaluation code, and model results associated with the paper.

jboye12/llm-probs dataset
Open GitHub

Repository containing the data and code for the LLM biased reinforcement learning experiments.

william-hayes/LLMs-biased-RL dataset
Open GitHub

The repository contains task implementations for the Iowa Gambling Task, Cambridge Gambling Task, and Wisconsin Card Sort Test, LLM integration utilities, oTree interface components, run scripts, configuration files, and links to analytical code and data.

ynulihao/LLM_vs_Human_Decision_Making benchmark
Open GitHub

Official code repository for the Rememberer agent and Reinforcement Learning with Experience Memory experiments.

OpenDFM/Rememberer
Open GitHub

Repository released by the authors for reproducing the Zero-shot-CoT experiments.

kojima-takeshi188/zero_shot_cot
Open GitHub

Google DeepMind repository containing code for prompt optimization, prompt evaluation, linear regression, and traveling-salesman experiments described in the paper.

google-deepmind/opro
Open GitHub

Repository for TraceLLM, the LLM-based synthetic microservice trace generator and related experimental code.

ldos-project/TraceLLM system
Open GitHub

Official source-code repository for the ICPI algorithm and experiments introduced in the paper.

ethanabrooks/icpi
Open GitHub

Repository for the paper's LLM-based anomaly detection implementation, including data-processing scripts, raw data archive, demo scripts for SFT, transfer learning, LoRA, online detection, catastrophic forgetting, and ICL.

PoSeiDon-Workflows/LLM_AD implementation
Open GitHub

A GitHub repository linked by the authors as a resource for recent work on LLMs for ML workflows.

t-harden/LLM4AutoML
Open GitHub

The repository contains the paper's source code in a Jupyter notebook and the RAG knowledge-base text used for the demonstration.

explainable-digital-twins/RAG-DDDAS system
Open GitHub

Repository for transformer and foundation models for financial time-series forecasting, including data folders, model implementations, scripts, result files, notebooks, and reproduction instructions.

UVA-MLSys/Financial-Time-Series dataset
Open GitHub

Fin-LLAMA: efficient fine-tuning of quantized LLMs for finance.

Bavest/fin-llama implementation
Open GitHub

Cornucopia-LLaMA-Fin-Chinese: Chinese finance-oriented LLaMA model referenced by the survey.

jerry1993-tech/Cornucopia-LLaMA-Fin-Chinese implementation
Open GitHub

Repository containing the entity-linking pipeline, relevant-document counting code, and data used for the paper's analysis.

nkandpa2/long_tail_knowledge
Open GitHub

Repository linked by the paper for released code; it contains Product-Key Memory Layer support and a PKM-layer notebook.

facebookresearch/XLM
Open GitHub

An up-to-date resource list for large multimodal agents associated with the survey paper.

jun0wanan/awesome-large-multimodal-agents
Open GitHub

Repository implementing the GANPO latent adversarial regularization framework introduced by the paper.

enyijiang/GANPO implementation
Open GitHub

Source-code repository for the TACO-RL algorithm, configurations, training pipeline, and CALVIN-based experiments introduced in the paper.

ErickRosete/tacorl
Open GitHub

Microsoft UniLM's LayoutLMv3 implementation, fine-tuning examples, and links to pre-trained and task-fine-tuned models.

microsoft/unilm implementation
Open GitHub

Repository backing the paper's companion website and editable table of reviewed studies, access modes, reproducibility attributes, comparison practices, and leaked datasets.

leak-llm/leak-llm.github.io
Open GitHub

The official code repository for experiments and evaluation associated with the Leaky Thoughts paper.

parameterlab/leaky_thoughts benchmark
Open GitHub

Repository containing the annotation backend, web frontend, model API, example notebooks, and explanation-supported training components for named entity recognition, relation extraction, and sentiment analysis.

INK-USC/LEAN-LIFE
Open GitHub

Repository linked by the paper as the available code for the conformal abstention and uncertainty-evaluation work.

sinatayebati/vlm-uncertainty benchmark
Open GitHub

Official implementation and experiments for Graph Neural Controlled Differential Equations.

WonderSeven/graph-neural-cdes framework
Open GitHub

Repository described as the official codebase for the ACL submission titled 'Learning Faster with Better Tokens: Parameter-Efficient Vocabulary Adaptation for Specialized Text Summarization'.

gb-kgp/VocabReplace-Then-Expand implementation
Open GitHub

Contains data resources and code for the two-stage training procedure, inference, and evaluation of fine-grained attributed generation.

LuckyyySTA/Fine-grained-Attribution dataset
Open GitHub

Repository for reproducing the FAIR paper, including scripts for inferring student mistakes, collecting teacher responses, training distilled student models, and testing accuracy.

zhuochunli/Learn-from-Committee framework
Open GitHub

Repository containing synthetic-data preparation, GRPO and SFT training code, experiment recipes, model evaluation scripts, and configurations used to reproduce the synthetic-to-real reasoning experiments.

kilian-group/phantom-reasoning
Open GitHub

Official implementation of the paper's TraceCodegen training framework, including preprocessing, execution, buffer-based self-sampling, training configurations, and inference.

microsoft/TraceCodegen
Open GitHub

Official implementation and pipeline repository for representation learning in time-domain high-energy astrophysics, including event-file representations, feature extraction, dimensionality reduction, clustering, datasets, encoders, and demonstration notebook.

StevenDillmann/ml-xraytransients-mnras implementation
Open GitHub

Repository titled 'BERTOps: Learning Representations on Logs for AIOps' containing data, results, source code, annotated datasets, dataset distributions, and scripts for preparing a pretrained ITOps-domain model.

BertOps/bertops dataset
Open GitHub

Official accompanying code repository for DemPref, including experiment code, raw simulation data, and Jupyter notebooks used to reproduce figures.

malayandi/DemPrefCode
Open GitHub

Code accompanying the paper's limited-expert-prediction learning-to-defer experiments.

ptrckhmmr/learning-to-defer-with-limited-expert-predictions implementation
Open GitHub

Repository containing code and data for the FILCO project, including context scoring, dataset preparation, model training, inference, and evaluation scripts.

zorazrw/filco implementation
Open GitHub

Official implementation of ProverGen, the framework that generates ProverQA, which is the primary benchmark used for PRoSFI training and evaluation.

opendatalab/ProverGen benchmark
Open GitHub

Repository containing code for the paper's Robot Language Model approach to grounded task planning.

dnandha/RobLM system
Open GitHub

Repository released by the authors for Learning to Self-Verify Makes Language Models Better Reasoners.

chenyuxin1999/Learning-to-Self-Verify
Open GitHub

Contains code for the supervised baseline, trained reward model, PPO fine-tuned policy, model card, and links to the released human-feedback and evaluation data.

openai/summarize-from-feedback dataset
Open GitHub

Official implementation and pretrained model weights for the CLIP method introduced and evaluated in the paper.

OpenAI/CLIP implementation
Open GitHub

Official repository for LeetCodeDataset, containing versioned data, data-processing code, and the eval_lcd evaluation toolkit.

newfacade/LeetCodeDataset dataset
Open GitHub

Repository explicitly linked by the paper for the open-sourced Lemur models and implementation resources.

OpenLemur/Lemur
Open GitHub

Repository for the ICML 2026 paper; it is the designated implementation and results release target for the benchmark.

CEA-LIST/less-precise-more-reliable-vlms
Open GitHub

Repository accompanying the paper with PRM800K labels, labeler instructions, grading logic, MATH train/test splits, and scored evaluation samples.

openai/prm800k
Open GitHub

Official experiment code for training LEVER verifiers, executing generated programs, and reproducing the paper's language-to-code evaluations.

niansong1996/lever
Open GitHub

CAIL2019 is used as an additional legal case similarity setting involving civil cases such as private lending, IP disputes, and maritime law.

alumik/cail2019 dataset
Open GitHub

Repository associated with the paper's rationale-augmented dialogue understanding experiments and released resources.

ShoRit/RATDIAL dataset
Open GitHub

Code, model-training scripts, refined UMLS relation data, architecture assets, and setup instructions for the DR.KNOWS system introduced and evaluated in the paper.

serenayj/DRKnows
Open GitHub

The repository contains training and evaluation code and pretrained-model support for Fusion-in-Decoder models.

facebookresearch/FiD
Open GitHub

Official source-code repository linked by the paper for the LeVo song-generation system.

tencent-ailab/songgeneration
Open GitHub

Repository containing the LitSearch dataset and code for constructing or evaluating the scientific literature retrieval benchmark.

princeton-nlp/LitSearch dataset
Open GitHub

Repository named IBM/live-api-bench with code for converting BIRD benchmark SQL queries into API call sequences and producing slot-filling and selection-style benchmark outputs.

IBM/live-api-bench dataset
Open GitHub

Repository for the LiveTradeBench platform/package used to run, monitor, and benchmark LLM-based trading agents across U.S. stock and Polymarket environments.

ulab-uiuc/live-trade-bench package
Open GitHub

Contains the LiveVectorLake Python implementation, chunk-level CDC components, Milvus and Delta Lake integrations, query engine, tests, generated test data, benchmark scripts, documentation, and architecture materials.

praj-tarun/LiveVectorLake dataset
Open GitHub

Official repository linked by the paper for access to the LLaMA model family and associated inference resources.

facebookresearch/llama
Open GitHub

Repository for the LlamaDuo LLMOps pipeline implementation.

deep-diver/llamaduo system
Open GitHub

Repository for the LLF-Bench environments, installation instructions, wrappers, and task implementations introduced by the paper.

microsoft/LLF-Bench benchmark
Open GitHub

Repository containing the nudging experimental framework, configuration for nudge types and models, and R analysis code for statistical tests and plots.

PapayaResearch/nudging benchmark
Open GitHub

Repository associated with the WORKS 2025 Flowcept Agent materials, including software, synthetic workflow, data, analysis code, query set, and prompts referenced for reproducibility.

flowcept/FlowceptAgent-WORKS25 system
Open GitHub

Flowcept code repository used as the provenance capture and observability foundation for the agent architecture and implementation.

ORNL/flowcept framework
Open GitHub

Repository for modernbert_predict_masked.

AnswerDotAI/ModernBERT
Open GitHub

Repository for medsam_inference.

bowang-lab/MedSAM
Open GitHub

Repository for flowmap_overfit_scene.

dcharatan/flowmap
Open GitHub

Repository for esm_fold_predict.

facebookresearch/esm
Open GitHub

Repository for stamp_extract_features and stamp_train_classification_model.

KatherLab/STAMP
Open GitHub

Public repository for TOOLMAKER code and TM-BENCH benchmark.

KatherLab/ToolMaker framework
Open GitHub

Repository for pathfinder_verify_biomarker.

LiangJunhao-THU/PathFinderCRC
Open GitHub

Repository for musk_extract_features.

lilab-stanford/MUSK
Open GitHub

Repository for conch_extract_features.

mahmoodlab/CONCH
Open GitHub

Repository for uni_extract_features.

mahmoodlab/UNI
Open GitHub

Repository for nnunet_train_model.

MIC-DKFZ/nnUNet
Open GitHub

Repository for medsss_generate.

pixas/MedSSS
Open GitHub

Repository for tabpfn_predict.

PriorLabs/TabPFN
Open GitHub

Repository for retfound_feature_vector.

rmaphoh/RETFound_MAE
Open GitHub

Repository for cytopus_db.

wallet-maker/cytopus
Open GitHub

Repository containing the cross-provider validation framework, SEC 10-K test corpus, synthetic financial database, 480 run traces, reproducibility manifests, and setup instructions for release v0.1.0.

ibm-client-engineering/output-drift-financial-llms dataset
Open GitHub

Repository containing a Python pipeline for extracting LLM uncertainty features, analysing and selecting features, balancing imbalanced data, training calibrated Ridge and XGBoost meta-models, tuning thresholds, and evaluating cost-aware uncertainty models.

ZEFR-INC/lpp-research framework
Open GitHub

Repository containing scripts and source code corresponding to the paper's main experiments, figures, and tables.

slhleosun/reasoning-trajectory implementation
Open GitHub

Official implementation of Chameleon, the plug-and-play compositional reasoning system reproduced and structurally analyzed in the paper.

lupantech/chameleon-llm
Open GitHub

Official repository containing the prompts, outputs, token-count materials, and supporting files for the paper's five experiments.

AKSW/AI-Tomorrow-2023-KG-ChatGPT-Experiments
Open GitHub

CAMEL is an open-source multi-agent framework for role-playing and agent collaboration.

camel-ai/camel framework
Open GitHub

Code repository for improving factuality and reasoning through multi-agent debate.

composable-models/llm_multiagent_debate framework
Open GitHub

MetaGPT is a multi-agent framework that models a software company using role assignments and SOP-style workflows.

FoundationAgents/MetaGPT framework
Open GitHub

AutoAgents generates different roles for GPTs to form a collaborative entity for complex tasks.

Link-AGI/AutoAgents framework
Open GitHub

Microsoft AutoGen, a framework for building multi-agent AI applications.

microsoft/autogen framework
Open GitHub

Code repository for Solo Performance Prompting / multi-persona self-collaboration.

MikeWangWZHL/Solo-Performance-Prompting framework
Open GitHub

AgentVerse provides task-solving and simulation frameworks for multiple LLM-based agents.

OpenBMB/AgentVerse framework
Open GitHub

ChatDev implements LLM-powered multi-agent collaboration for software development.

OpenBMB/ChatDev framework
Open GitHub

Repository connected to AI Scientist-generated papers reported as having passed peer review at an ICLR workshop.

SakanaAI/AI-Scientist-ICLR2025-Workshop-Experiment
Open GitHub

Code repository for MAD, a multi-agent debate framework using large language models.

Skytliang/Multi-Agents-Debate framework
Open GitHub

A curated collection of resources associated with the survey, described by the authors as containing over 200 related papers on agent hallucinations.

ASCII-LAB/Awesome-Agent-Hallucinations
Open GitHub

Repository containing code, scripts, data folders, requirements, and reproduction instructions for LLM-DSE: Searching Accelerator Parameters with LLM Agents.

Nozidoali/LLM-DSE framework
Open GitHub

Repository reported by the paper as containing the source code and data for the LLM-BLM study.

youngandbin/LLM-BLM framework
Open GitHub

Repository for reproducing LLM-Explorer experiments and implementation details.

tsinghua-fib-lab/LLM-Explorer framework
Open GitHub

Official repository linked from the paper's project page; it contains the Chat with NeRF/LLM-Grounder implementation, interactive demo code, and notebooks for reproducing the paper's results.

sled-group/chat-with-nerf
Open GitHub

Official repository for LLM-Planner. It provides an hlp directory with the high-level prompt generator and kNN dataset from the paper, plus an e2e directory with an end-to-end agent using LLM-Planner.

OSU-NLP-Group/LLM-Planner
Open GitHub

Repository for the paper's LLM-REVal framework, including simulation components for LLM-driven research and review workflows.

PlusLabNLP/LLM-REVal benchmark
Open GitHub

Repository containing the LLM-SR code and the generated Oscillation 1, Oscillation 2, and E. coli Growth benchmark datasets.

deep-symbolic-mathematics/LLM-SR implementation
Open GitHub

Repository for BIG-Bench Mistake, including the dataset, prompting materials, annotation guidelines, and released code referenced throughout the paper.

WHGTyen/BIG-Bench-Mistake dataset
Open GitHub

Contains the DELEGATE-52 relay runners, direct and agentic model wrappers, prompts, and 52 domain-specific parsers and evaluators.

microsoft/DELEGATE52 dataset
Open GitHub

Official implementation of the Chronos time-series foundation model used for zero-shot inference and sequential fine-tuning.

amazon-science/chronos-forecasting framework
Open GitHub

Configuration files for the CNN-Transformer statistical-arbitrage benchmark replicated in the paper.

gregzanotti/dlsa-public benchmark
Open GitHub

Repository directory containing the IPCA, PCA, and Fama-French residual-return datasets used in the paper's backtests.

gregzanotti/dlsa-public dataset
Open GitHub

Repository for the extensible Judge-Bench collection, shared data schema, prompts, and evaluation code used to compare LLM judgments with human annotations.

dmg-illc/JUDGE-BENCH benchmark
Open GitHub

Official code repository for the paper, containing evaluation scripts, configuration files, a CSV of evaluation results, notebooks, figures, and a modified Lingua training submodule.

brendel-group/llm-line implementation
Open GitHub

Official repository containing the LLMSecEval dataset, code-generation application, and CodeQL-based security-analysis materials introduced and demonstrated in the paper.

tuhh-softsec/LLMSecEval dataset
Open GitHub

Repository for the Needle-In-A-Haystack long-context retrieval test adapted in the paper's fixed- and random-needle experiments.

gkamradt/needle-in-a-haystack benchmark
Open GitHub

Official repository for the LLoCO framework and experiments.

jeffreysijuntan/lloco framework
Open GitHub

Official LongBench repository used for the paper's SingleDoc, MultiDoc, and summarization evaluation; baseline numbers are taken from this repository.

THUDM/LongBench dataset
Open GitHub

Google DeepMind repository containing LMAct environments, expert agents, prompt construction, evaluation code, experiment configuration, and instructions for downloading expert demonstrations.

google-deepmind/lm_act benchmark
Open GitHub

Repository for Local-Splitter, an MCP-compatible and OpenAI-compatible outbound LLM request shim that uses a local small model as a triage layer to reduce cloud token usage.

jayluxferro/local-splitter system
Open GitHub

GitHub repository linked by the paper for LocalEval resources; the repository page inspected stated that available resources were being prepared.

tsinghua-fib-lab/LocalEval dataset
Open GitHub

Public repository for LoCoBench-Agent, including the benchmark framework, evaluation setup, data-download instructions, and metrics for long-context software engineering agent evaluation.

SalesforceAIResearch/LoCoBench-Agent dataset
Open GitHub

Repository containing LogicBench data, evaluation code, and supporting reasoning-chain materials.

Mihir3009/LogicBench dataset
Open GitHub

Public-facing repository for the paper's KD experiments, configuration files, training code, plotting scripts, and geometric feature analysis utilities.

Thegolfingocto/KD_wo_CE benchmark
Open GitHub

Official repository for the LongBench benchmark, datasets, and evaluation code introduced and evaluated in the paper.

THUDM/LongBench dataset
Open GitHub

Repository for the LongMemEval benchmark, evaluation code, history-construction tools, and memory-system experiments.

xiaowu0162/LongMemEval dataset
Open GitHub

Repository containing LongReasonArena benchmark data, input-generation utilities, inference code, and evaluation scripts.

LongReasonArena/LongReasonArena dataset
Open GitHub

Repository containing code, data, and prompts for evaluating LLMs as path planners, including the artifact associated with 'Look Further Ahead: Testing the Limits of GPT-4 in Path Planning'.

MohamedAghzal/llms-as-path-planners benchmark
Open GitHub

Public implementation of the Loquetier virtualized multi-LoRA framework, including kernel source, runtime code, examples, tests, and licenses.

NJUDeepEngine/Loquetier
Open GitHub

Repository containing preprocessing code, prompts, configurations, dataset splits, response extraction, and LoRAX load-testing scripts used in the report.

predibase/lora_bakeoff benchmark
Open GitHub

Open-source multi-LoRA inference server using shared base weights, dynamic adapter loading, cross-adapter batching, and adapter-weight caching.

predibase/lorax
Open GitHub

Official repository containing the loralib Python package, LoRA integration examples, model checkpoints, and reproduction code for RoBERTa, DeBERTa, and GPT-2 experiments.

microsoft/LoRA
Open GitHub

Repository containing the data, curve-fitting code, theory code, and notebooks used to generate the paper's empirical and appendix figures.

KempnerInstitute/loss-to-loss-notebooks
Open GitHub

Training code used for the loss-to-loss prediction experiments, including sweep configuration and OLMo-based model training infrastructure.

KempnerInstitute/loss-to-loss-olmo
Open GitHub

Official repository containing the paper's QA and key-value datasets, prompt construction code, data-generation scripts, tests, and experiment instructions.

nelson-liu/lost-in-the-middle
Open GitHub

Repository for the LoTR method introduced and experimentally evaluated in the paper.

skolai/lotr
Open GitHub

Repository for LPS-BENCH containing benchmark examples, mock tools, evaluators, prompt templates, schemas, scripts, and the multi-agent case synthesis pipeline.

tychenn/LPS-Bench dataset
Open GitHub

Repository released for the Lynx hallucination-detection models and HaluBench research artifacts, including code, training data, and model generations.

patronus-ai/Lynx-hallucination-detection
Open GitHub

Repository released by the authors for the embedding model family and the paper's model/code/data resources.

FlagOpen/FlagEmbedding
Open GitHub

Official repository containing the M3Bench dataset organization, PyBullet and Isaac Sim evaluation workflow, TongVerse setup, task configurations, and instructions for evaluating generated trajectories.

TooSchoolForCool/M3Bench benchmark
Open GitHub

Repository for the general-purpose MadEvolve code-evolution framework with MAP-Elites, island populations, multiple LLM providers, evaluation backends, and analysis tools.

tianyi-stack/MadEvolve framework
Open GitHub

Repository for the Self-Supervised Audio Spectrogram Transformer used as the architectural and empirical baseline for MAE-AST.

YuanGongND/ssast implementation
Open GitHub

Repository for MaGNet containing Python implementation files for MAGE, hypergraph modules, 2D attention modules, training, datasets/model-weight download links, and backtesting.

PeilinTime/MaGNet dataset
Open GitHub

Official code repository for Maieutic Prompting, including tree generation, inference code, benchmark data files, and pre-generated maieutic trees.

jaehunjung1/Maieutic-Prompting implementation
Open GitHub

Repository for the MALLM-GAN method, including the main model code, an Adult-dataset notebook, baseline pipeline, sampling utilities, evaluation utilities, and dependency specifications.

yling1105/MALLM-GAN implementation
Open GitHub

Repository linked by the paper for Mamba model code and pretrained checkpoints.

state-spaces/mamba
Open GitHub

Repository containing the ManipulaTHOR environment wrapper, ArmPointNav task and samplers, APND dataset instructions, baseline configurations, actor-critic models, evaluation scripts, and pretrained-model guidance.

allenai/manipulathor dataset
Open GitHub

Repository containing the authors' simulator and market-making agents for scaled beta policy experiments.

JJJerome/rl4mm framework
Open GitHub

Repository for the FOREC/Cross-Market Product Recommendation baseline used for model parameters and comparisons.

hamedrab/FOREC implementation
Open GitHub

Repository linked by the paper for the efficient cross-market recommendation implementation associated with the proposed market-aware models.

samarthbhargav/efficient-xmrec implementation
Open GitHub

Repository released by the authors for Markov Chain of Thought code and associated research artifacts.

james-yw/Markov-Chain-of-Thought dataset
Open GitHub

Repository for MASEval, the framework-agnostic multi-agent system evaluation library introduced and evaluated in the paper.

parameterlab/MASEval framework
Open GitHub

PyTorch/GPU re-implementation of MAE with visualization, pre-training, linear-probing and fine-tuning code, plus pretrained checkpoints used for the paper's reported models.

facebookresearch/mae implementation
Open GitHub

Code repository for the Point-MAE framework introduced and evaluated in the paper.

Pang-Yatian/Point-MAE implementation
Open GitHub

Repository hosting Audio-MAE code and pretrained models for masked spectrogram autoencoding and downstream audio tasks.

facebookresearch/AudioMAE implementation
Open GitHub

Provides code and model assets for MASS pre-training and fine-tuning, including unsupervised and supervised NMT, text summarization, and conversational response generation.

microsoft/MASS implementation
Open GitHub

Repository for the MathCoder family, including inference and evaluation code and links to the MathCodeInstruct dataset and released model checkpoints.

mathllm/MathCoder
Open GitHub

Hosts the MathPile project documentation and source-processing code associated with the corpus introduced and evaluated in the paper.

GAIR-NLP/MathPile dataset
Open GitHub

Official MathVista repository containing benchmark data access instructions, evaluation scripts, model-output templates, leaderboard submission resources, and code associated with the paper.

lupantech/MathVista dataset
Open GitHub

Official repository for the MatPlotAgent framework and MatPlotBench benchmark introduced and evaluated in the paper.

thunlp/MatPlotAgent benchmark
Open GitHub

Author-owned repository with ChartQA, FigureQA, PlotQA, and test-file directories for the paper's complex chart-question-answering data.

ChangGuiyong/mChartQA dataset
Open GitHub

Official code repository for running the MDAgents framework and reproducing benchmark experiments across supported medical datasets and LLM backends.

mitmedialab/MDAgents
Open GitHub

Official repository for the mDPO method, including an mDPO trainer, Bunny training code, and data links.

luka-group/mDPO implementation
Open GitHub

Sibyl System repository included in the selected AutoGen application sample.

Ag2S1/Sibyl-System system
Open GitHub

AutoTx repository for planning and executing on-chain transactions.

agentcoinorg/AutoTx
Open GitHub

GPT-Academic repository for LLM-assisted academic reading, writing, translation, and code/project analysis workflows.

binary-husky/gpt_academic
Open GitHub

Composio platform/repository evaluated as a flexible agent application or platform with multiple autonomy-related configurations.

ComposioHQ/composio
Open GitHub

h2oGPT repository for private local GPT-style chat and document interaction.

h2oai/h2ogpt
Open GitHub

GraphRag_Ollama repository combining AutoGen, GraphRAG, Ollama, and related tooling.

karthikvenkatesan-eaton/Autogen_GraphRAG_Ollama
Open GitHub

Langflow platform for building and deploying AI-powered agents and workflows.

langflow-ai/langflow
Open GitHub

Letta platform for stateful agents with advanced memory.

letta-ai/letta
Open GitHub

AutoGen open-source framework for building AI agent systems using language models, multi-agent conversations, and tool use.

microsoft/autogen framework
Open GitHub

AutoGen Studio application within the AutoGen repository.

microsoft/autogen system
Open GitHub

Dream Team repository for building a team of AI agents with AutoGen.

yanivvak/dream-team
Open GitHub

Repository containing the paper's self-ask implementation/demo and released datasets, including materials for Compositional Celebrities and Bamboogle.

ofirpress/self-ask
Open GitHub

GitHub source for the multiple-choice Truthful-QA variant used in the model-level ranking experiment.

manyoso/haltt4llm dataset
Open GitHub

CHALE repository used as a hallucination-evaluation dataset with non-hallucinated, half-hallucinated, and hallucinated answer categories.

weijiaheng/CHALE dataset
Open GitHub

Official repository for Measuring Massive Multitask Language Understanding, containing evaluation and calibration code and links to download the test.

hendrycks/test dataset
Open GitHub

Search/RAG infrastructure repository.

devflowinc/trieve
Open GitHub

Large language model training framework repository.

EleutherAI/gpt-neox
Open GitHub

NLP framework repository.

flairNLP/flair
Open GitHub

Glasgow Haskell Compiler repository.

ghc/ghc
Open GitHub

Haskell Cabal build/package repository.

haskell/cabal
Open GitHub

Transformer model library repository.

huggingface/transformers
Open GitHub

Property-based testing library repository.

HypothesisWorks/hypothesis
Open GitHub

JavaScript DOM implementation repository.

jsdom/jsdom
Open GitHub

Python spreadsheet/data-analysis tool repository.

mito-ds/mito
Open GitHub

Machine-learning library repository.

scikit-learn/scikit-learn
Open GitHub

JavaScript standard library monorepo.

stdlib-js/stdlib
Open GitHub

Official MechAgents repository containing Colab notebooks for the two-agent elasticity and hyperelasticity workflows, the multi-agent linear-elasticity workflow, and the direct two-agent versus multi-agent comparison.

lamm-mit/MechAgents
Open GitHub

Repository associated with the paper's Time Series Transformer mechanistic interpretability experiments.

mathiisk/TST-Mechanistic-Interpretability implementation
Open GitHub

Repository for the MedBayes-Lite clinical uncertainty governance layer and associated experimental implementation.

eliashossain001/medbayes-lite framework
Open GitHub

Official repository for the MegaMath dataset project, including released web and code pipeline materials and instructions for accessing dataset variants.

LLM360/MegaMath dataset
Open GitHub

Public codebase for the MELT on-device LLM evaluation system, including PhoneLab and JetsonLab infrastructure, framework integrations, model conversion/evaluation code, prompt analysis, and result parsing.

brave-experiments/MELT-public
Open GitHub

Repository containing the official Python implementation of MEME, including code for argument extraction, mode identification/alignment, and training scripts; the README notes that the dataset used in the project is private due to compliance requirements.

gta0804/MEME framework
Open GitHub

Repository for the MeMemo browser-based HNSW retrieval toolkit, documentation, and RAG Playground example application.

poloclub/mememo framework
Open GitHub

Canonical repository for the MemGPT system, now named Letta, implementing stateful LLM agents with persistent memory and related tooling.

letta-ai/letta
Open GitHub

Repository for MemR3, the memory retrieval via reflective reasoning controller.

Leagein/memr3 framework
Open GitHub

Official open-source implementation of the MemTools framework introduced and evaluated in the paper.

JJJAYYYZhao/MemTools-public
Open GitHub

Repository providing data and code for reproducing analyses, with folders for one-dimensional generated regression, two-armed bandit, real-world regression, and an MMLU benchmark.

juliancodaforno/meta-in-context-learning benchmark
Open GitHub

Official project repository for the MetaGPT framework introduced and evaluated in the paper.

geekan/MetaGPT implementation
Open GitHub

Repository containing the MetaOptimize implementation, package files, example usage, and experiment code for CIFAR10, ImageNet, TinyStories, and continual CIFAR100 experiments.

sabersalehk/MetaOptimize framework
Open GitHub

Official MetaTool repository containing ToolE data, tool descriptions and embeddings, scenario lists, prompt templates, model-generation scripts, and evaluation code.

HowieHwong/MetaTool benchmark
Open GitHub

Public implementation companion containing the intent-to-specification pre-experiment stage, SevenNet-based experiment utilities and MCP server scripts, and discussion, revision, CLI, and Streamlit interface components.

IMMS-Ewha/MIND
Open GitHub

Official repository for the Mind2Web dataset, benchmark processing and evaluation resources, and MindAct fine-tuning and model code.

OSU-NLP-Group/Mind2Web dataset
Open GitHub

Contains the released evaluation code, Graphviz-based synthetic mind-map generator, crawled and synthetic annotation preparation tools, and links to the benchmark datasets.

MiaSanLei/MindBench benchmark
Open GitHub

Repository linked by the paper for the MindWatcher agent framework, models, benchmark resources, and related implementation artifacts.

TIMMY-CHAN/MindWatcher framework
Open GitHub

Repository for MiniCheck code, model usage, synthetic data generation code, benchmark evaluation demo, and links to LLM-AggreFact and Hugging Face model resources.

Liyan06/MiniCheck benchmark
Open GitHub

Repository for the Minions communication protocol enabling small on-device models to collaborate with frontier cloud models.

HazyResearch/minions system
Open GitHub

Official repository for reproducing MINT evaluation, benchmark data, prompts, and supporting code.

xingyaoww/mint-bench dataset
Open GitHub

Code repository for the MISS generative medical VQA framework and its training approach.

TIMMY-CHAN/MISS implementation
Open GitHub

Repository containing the paper's mKG-RAG implementation and associated resources.

xandery-geek/mKG-RAG framework
Open GitHub

MedicalZooPytorch repository used in the benchmark's OOD training set.

black0017/MedicalZooPytorch
Open GitHub

Text classification repository listed as TCL in the benchmark's OOD training repositories.

brightmart/text_classification
Open GitHub

DeepFloyd IF repository used for multi-modal image tasks.

deep-floyd/if
Open GitHub

Deep Graph Library repository used for graph-model tasks.

dmlc/dgl framework
Open GitHub

PyTorch-GAN repository used for image-GAN tasks.

eriklindernoren/PyTorch-GAN
Open GitHub

ESM repository used for protein/biomedical tasks.

facebookresearch/esm
Open GitHub

Public ML-Bench code and benchmark resources released by the paper.

gersteinlab/ML-bench dataset
Open GitHub

BERT repository used for ML-Bench tasks.

google-research/bert
Open GitHub

PyTorch Image Models repository used as an OOD evaluation repository.

huggingface/pytorch-image-models
Open GitHub

Grounded Segment Anything repository used as an OOD evaluation repository.

IDEA-Research/Grounded-Segment-Anything
Open GitHub

Muzic repository used for music/audio tasks.

microsoft/muzic
Open GitHub

OpenCLIP repository used for multi-modal tasks.

mlfoundations/open_clip
Open GitHub

vid2vid repository used for video tasks.

NVIDIA/vid2vid
Open GitHub

OpenDevin agent framework evaluated in ML-Agent-Bench with GPT-4o, GPT-4, and GPT-3.5.

OpenDevin/OpenDevin system
Open GitHub

LAVIS repository used for multi-modal tasks.

salesforce/lavis framework
Open GitHub

Stable Diffusion repository used in the benchmark's OOD training set.

Stability-AI/stablediffusion
Open GitHub

Tensor2Tensor repository used in the benchmark's OOD training set.

tensorflow/tensor2tensor
Open GitHub

Time-Series-Library repository used for time-series tasks.

thuml/Time-Series-Library
Open GitHub

Learning3D repository used for 3D vision tasks.

vinits5/learning3d
Open GitHub

External-Attention-pytorch repository used for attention-use tasks.

xmu-xiaoma666/External-Attention-pytorch
Open GitHub

Official code repository for MLAgentBench, including benchmark tasks, environment, agent implementations, runners, and evaluation scripts.

snap-stanford/MLAgentBench benchmark
Open GitHub

Repository for the benchmark and evaluation resources introduced by the paper.

chchenhui/mlrbench dataset
Open GitHub

Repository for mMARCO with translation scripts, dataset access information, retrieval instructions, and released fine-tuned model references.

unicamp-dl/mMARCO dataset
Open GitHub

Official MMDT source repository with scripts and framework components for evaluating text-to-image and image-to-text models across the six trustworthiness perspectives.

AI-secure/MMDT
Open GitHub

Official repository containing inference and evaluation code, setup instructions, and links to the released dataset and leaderboard.

Alpha-Innovator/MME-Reasoning dataset
Open GitHub

Official repository for the MMIR-TCM framework and its associated MedTCM and TDEU artifacts.

jw-chae/MMIR-TCM
Open GitHub

Repository for the MMReason benchmark, with citation and evaluation instructions and integration through VLMEvalKit.

HJYao00/MMReason dataset
Open GitHub

Official XiaoMi repository for Mobile-Bench, including Appium/emulator setup materials, API/UI agent code folders, data/result directories, prompts, requirements, and README instructions.

XiaoMi/MobileBench dataset
Open GitHub

Repository for MobileAgentBench, an automated benchmark for mobile LLM agents with a Python library and default tasks using SimpleMobileTools apps.

MobileAgentBench/mobile-agent-bench benchmark
Open GitHub

Repository named Modalitites-Impact-In-ML under the belgats GitHub account. It was public but empty when accessed.

belgats/Modalitites-Impact-In-ML implementation
Open GitHub

Public Python repository containing MODE ingestion and inference code, clustering and centroid-routing components, benchmark scripts, evaluation data and logs, tests, and documentation.

rahulanand1103/mode framework
Open GitHub

Official Model Context Protocol server collection used to identify official and community MCP integrations.

modelcontextprotocol/servers
Open GitHub

Replication package released by the authors for the MCP server empirical study.

SAILResearch/replication-25-mcp-server-empirical-study dataset
Open GitHub

Public repository containing the paper's collected landscape data and implementation examples.

security-pride/MCP_Landscape dataset
Open GitHub

Python/TensorFlow implementation of Robust Log-Optimal Strategy with Reinforcement Learning discussed as a policy-based portfolio optimization approach.

fxy96/Robust-Log-Optimal-Strategy-with-Reinforcement-Learning framework
Open GitHub

Python/TensorFlow implementation of the PGPortfolio or EIIE-style portfolio reinforcement learning framework discussed in the survey.

ZhengyaoJiang/PGPortfolio framework
Open GitHub

GitHub repository for the SYMBA crypto multi-agent reinforcement learning market simulator.

johannlussange/symba_crypto framework
Open GitHub

Official implementation repository for the ModuleFormer architecture, with inference code and links to released MoLM model weights.

IBM/ModuleFormer implementation
Open GitHub

The Megatron-DeepSpeed repository is explicitly linked as the DeepSpeed implementation used as the principal large-scale MoE comparison baseline.

microsoft/Megatron-DeepSpeed
Open GitHub

PaddleFleetX is the open-source architecture on which MoESys is implemented and is benchmarked against Megatron-LM in the platform comparison.

PaddlePaddle/PaddleFleetX
Open GitHub

Repository for the MonkeyOCR model family, including inference code, deployment instructions, model links, benchmark materials, and a dated announcement for the MonkeyOCR v1.5 technical report.

Yuliang-Liu/MonkeyOCR implementation
Open GitHub

Repository containing the paper's fine-tuning, inference, plotting, and experiment configuration code for intrinsic moral reward training of LLM agents.

liza-tennant/LLM_morality
Open GitHub

Repository for the GroundMoRe dataset workflow and MoRA training, fine-tuning, checkpoint conversion, and evaluation code.

dengandong/GroundMoRe
Open GitHub

Repository for the mPLUG-DocOwl model family and the paper's released code, models, training data, and evaluation resources.

X-PLUG/mPLUG-DocOwl implementation
Open GitHub

Synthetic train, development, and test files used for the paper's arithmetic argument-extraction experiments.

AI21Labs/MRKL_synthetic_data dataset
Open GitHub

TensorFlow Datasets implementation for distributing the MT-Opt robot-episode and success-detector datasets associated with the paper.

tensorflow/datasets dataset
Open GitHub

Auto-GPT is discussed as an autonomous AI application whose main agent, plugins, memory, file access, internet access, code execution, image generation, and oracle-like functions can be represented within the proposed multi-agent graph framework.

Significant-Gravitas/Auto-GPT framework
Open GitHub

BabyAGI is discussed as an AI agent system with task creation, prioritisation, and execution chains that can be modelled as interconnected agents plus a vector-database plugin.

yoheinakajima/babyagi system
Open GitHub

Official source-code repository for the paper, including scalar and two-dimensional consensus experiments, topology configuration, plotting utilities, and experiment execution code.

WindyLab/ConsensusLLM-code
Open GitHub

Official data release containing generated experiment outputs for the agent-count, temperature, and personality analyses.

WindyLab/ConsensusLLM-code
Open GitHub

Repository associated with the experiment in which a fine-tuned Llama 2 7B Chat model attempts to manipulate an LLM overseer or reward model, with RL increasing jailbreak attempts.

AlexMeinke/fooling-the-overseer implementation
Open GitHub

Repository associated with the experiment that repeatedly rewrites news articles using LLMs and evaluates degradation in factual accuracy.

qfeuilla/DistordedNews implementation
Open GitHub

Repository associated with the experiment testing whether specialized LLM driving agents fine-tuned on different traffic conventions fail to coordinate when yielding to an emergency vehicle.

SUMEETRM/driving_llms implementation
Open GitHub

Official X-VLM repository containing PyTorch pre-training and fine-tuning code, task configurations, data-format examples, and links to pre-trained and fine-tuned checkpoints.

zengyan-97/X-VLM
Open GitHub

Official MH-MoE implementation based on TorchScale and fairseq, including setup and pretraining scripts.

yushuiwx/MH-MoE implementation
Open GitHub

Repository created to support the survey by collecting and categorizing relevant research papers, datasets, application scenarios, framework figures, and future-direction materials on value alignment in agentic AI systems.

Wei-ZENG1020/Value-Alignment-Agentic-AI-Papers-Survey-Taxonomy
Open GitHub

Repository containing the Multi-LogiEval data and associated evaluation or reasoning-chain artifacts.

Mihir3009/Multi-LogiEval dataset
Open GitHub

GitHub repository made available by the authors for MMTB.

yupeijei1997/MMTB dataset
Open GitHub

Source code and experiment framework for the paper, including MGDA-UB/min-norm solvers and MultiMNIST-related code/data.

isl-org/MultiObjectiveOptimization
Open GitHub

Repository containing the official implementation of Agentic Predictor for multi-view performance prediction in LLM-based agentic workflows.

DeepAuto-AI/agentic-predictor system
Open GitHub

Public repository for the MARBLE framework and MultiAgentBench code and data used to develop, test, and evaluate LLM-based multi-agent systems.

MultiagentBench/MARBLE dataset
Open GitHub

Repository identified by the paper as containing the project, data, Python code, ontology model, mapping rules, hyperparameter tuning code for ComplEx, TransE, and DistMult, and code to train and predict using TransE.

durgeshnandini/Multidimensional-Knowledge-Graph-Embeddings-for-International-Trade-Flow-Analysis dataset
Open GitHub

Repository for reproducing baseline generation, citation generation, MuRGAt-Score evaluation, and program-aided generation experiments.

meetdavidwan/murgat benchmark
Open GitHub

Official code and model workflow for preprocessing SIMMC 2.0 images, ITM and BTM pretraining, task-specific training, and challenge evaluation.

rungjoo/simmc2.0
Open GitHub

Repository containing MedTok training and inference code, dataset preparation folders, EHR and medical-QA tutorials, environment configuration, and links to released token embeddings.

mims-harvard/MedTok
Open GitHub

A constantly updated paperlist related to the survey topic of multimodality representation learning.

marslanm/multimodality-representation-learning
Open GitHub

Public repository containing modules that form MRT and utility code used for data piping in the paper's experiments.

EgoPer/Multiple-Resolution-Tokenization implementation
Open GitHub

Toolkit and repository for creating, sharing, and materializing the natural-language prompt templates used to build P3 and train T0.

bigscience-workshop/promptsource
Open GitHub

Code and instructions for reproducing T0 training, evaluation, inference, and ablation checkpoints.

bigscience-workshop/t-zero
Open GitHub

Official author repository containing MuSiQue data download utilities, evaluation scripts, released predictions, trained-model utilities, and code/configurations for the paper's experiments.

stonybrooknlp/musique dataset
Open GitHub

LASSO is a platform for scalable software code analysis and observation, combining dynamic and static program analysis, code search, N-version assessment, automated test generation, software experimentation, and benchmarking.

SoftwareObservatorium/lasso framework
Open GitHub

Source code used for the N^2M^2 paper, including training and evaluation scripts, robot environments, task wrappers, Gazebo assets, and instructions for PR2, HSR, and TIAGo experiments.

robot-learning-freiburg/mobile-rl
Open GitHub

Supplemental repository containing accepted submissions, the call for papers, nanopublication files, the questionnaire, and visualization materials for the field study.

LaraHack/formalization_papers_supplemental
Open GitHub

Repository containing the nanopublication collection and graph assets used to analyze and visualize the formalization-paper special-issue workflow.

LaraHack/fpsi_analytics
Open GitHub

Repository for Nanobench, the template-driven interface used by participants to create formalizations, class definitions, submissions, reviews, responses, updates, and decisions as nanopublications.

peta-pico/nanobench
Open GitHub

Repository for Tapas, the generic triple-store interface used to run template-based SPARQL queries and display submission and review overviews.

peta-pico/tapas
Open GitHub

Repository providing an open-source NRT-style training recipe built on top of verl, including scripts and configuration for NRT training.

sharkwyf/native-reasoning-models framework
Open GitHub

CodRep dataset/competition artifact for code refinement and defect detection.

ASSERT-KTH/CodRep dataset
Open GitHub

BigCloneBench dataset for clone detection and related code-understanding tasks.

clonebench/BigCloneBench dataset
Open GitHub

Code Contests dataset for code generation from competitive-programming problems.

deepmind/code_contests dataset
Open GitHub

InCoder repository for generative code models.

dpfried/incoder implementation
Open GitHub

Description2Code dataset for mapping descriptions to code.

ethancaballero/description2code dataset
Open GitHub

CodeSearchNet dataset covering multiple programming languages for code search and code-language tasks.

github/CodeSearchNet dataset
Open GitHub

APPS benchmark for code generation from programming problems.

hendrycks/apps benchmark
Open GitHub

Project CodeNet dataset for program understanding, generation, and refinement tasks.

IBM/Project_CodeNet dataset
Open GitHub

Copilot for Xcode source editor extension integrating GitHub Copilot and ChatGPT-style functions into Xcode.

intitni/CopilotForXcode system
Open GitHub

CodeXGLUE benchmark for multiple code intelligence tasks.

microsoft/CodeXGLUE dataset
Open GitHub

PyCodeGPT/CERT-related repository for Python code generation.

microsoft/PyCodeGPT implementation
Open GitHub

xCodeEval/ExecEval repository for multilingual code evaluation tasks.

ntunlp/xCodeEval benchmark
Open GitHub

HumanEval benchmark for evaluating code-generation models.

openai/human-eval benchmark
Open GitHub

WikiSQL dataset for SQL generation/summarization-style tasks in the paper's table.

salesforce/WikiSQL dataset
Open GitHub

CONCODE dataset for mapping natural language and program context to code.

sriniiyer/concode dataset
Open GitHub

Code-LMs repository associated with PolyCoder and code-language-model research.

VHellendoorn/Code-LMs implementation
Open GitHub

Official implementation repository for the expanded Natural Language Reinforcement Learning project, including shared NLRL libraries and later Maze, Breakthrough, and Tic-Tac-Toe experiments.

waterhorse1/Natural-language-RL
Open GitHub

Python code for collecting LLM ranking data, training the scoring model, and training policies with direct-score or potential-difference rewards.

sy-shi/RLAIF_ScoreDiff implementation
Open GitHub

GroundHog is the authors' Theano-based recurrent neural network framework containing the neural machine translation implementation used for the paper.

lisa-groundhog/GroundHog
Open GitHub

Repository released by the authors for Multimodal-ZeroShotTM, Multimodal-Contrast, and the multimodal topic modeling experiments.

gonzalezf/multimodal_neural_topic_modeling benchmark
Open GitHub

Repository for NNHedge, containing components for simulated instruments, neural hedging models, data loading, training, and assessment.

guijinSON/NNHedge framework
Open GitHub

Author-linked reference implementation of NSR, including SVD compression, TaskKnowledgeBank logic, allocation policies, dataset builders, experiment runners, and visualization code.

bhyoon-me/nsr
Open GitHub

Repository containing code for persona vector generation, the user-study interface, and user-study analysis associated with the neural transparency paper.

mitmedialab/neural-transparency system
Open GitHub

Repository containing code, data folders, source files, scripts, requirements, and usage instructions for running CaRing and evaluating ProofWriter, GSM8K, and PrOntoQA experiments.

DAMO-NLP-SG/CaRing framework
Open GitHub

Repository for the Next-Generation LLM for UAV system, including the main application, short-, medium-, and long-range route-planning utilities, path-planning utilities, control-platform utilities, examples, and integrated data folders.

liangqiyuan/NeLV framework
Open GitHub

Repository cited as the source of the Azure LLM inference production trace used to construct a dynamic request-rate workload.

Azure/AzurePublicDataset dataset
Open GitHub

The repository contains classifier, data/query, notebook, red_lm, and target_lm components plus training code; its README states that supervised-learning code is on the master branch and RL training is on the feature/ppo branch.

anugyas/NLUProject
Open GitHub

The repository describes NormCode as a language for auditable multi-step AI workflows, lists core components such as infra, canvas_app, cli_orchestrator.py, documentation, and examples, and links to the arXiv paper.

lys5588/Normcode-paper framework
Open GitHub

The fairseq NormFormer example provides architecture flags and causal-language-model training commands corresponding to the paper.

facebookresearch/fairseq implementation
Open GitHub

Official repository for NOSA and NOSI, including the attention implementation, NOSI package, environment instructions, benchmark scripts, and links to released NOSA-1B, NOSA-3B, and NOSA-8B models.

thunlp/NOSA
Open GitHub

Official PyTorch implementation of the paper's layer-wise expert pruning, progressive pruning, and dynamic expert skipping methods.

Lucky-Lance/Expert_Sparsity
Open GitHub

Repository containing code and configuration for training and inference of Qwen3 8B, DeepSeek LLM 7B, and LLaMA3 8B Instruct on financial sentiment datasets, including preprocessing, model configuration, training, inference, and evaluation.

NLPforFinance/fine-tuning-of-lightweight-large-language-models implementation
Open GitHub

Repository containing demonstrations and scenarios for indirect prompt-injection attacks, including synthetic GPT-based applications and attack examples discussed in the paper.

greshake/llm-security
Open GitHub

Open-source GPT-Researcher framework used as the fixed deep research execution pipeline and unpruned baseline on which pruning variants are implemented.

assafelovic/gpt-researcher
Open GitHub

Repository whose README cites the NumHTML AAAI 2022 paper and states that code and data used for the paper are provided.

YangLinyi/HTML-Hierarchical-Transformer-based-Multi-task-Learning-for-Volatility-Prediction dataset
Open GitHub

Official repository for O-Researcher, including model serving, web-search and crawl-page tools, inference code, deployment scripts, and links to released SFT/RL models and datasets.

OPPO-PersonalAI/O-Researcher
Open GitHub

Official repository for the Oasis data-curation and assessment system introduced by the paper.

tongzhou21/Oasis
Open GitHub

Repository containing the paper's Off-Policy Corrected Reward Modeling implementation and the code and resulting data for filtering the short Alpaca-Farm setting.

JohannesAck/OffPolicyCorrectedRewardModeling implementation
Open GitHub

AutoGen repository into which the paper states AgentOptimizer was integrated.

microsoft/autogen
Open GitHub

Official repository for the OLMoE paper, linking the released model, data, code, logs, configurations, and related resources.

allenai/OLMoE
Open GitHub

Repository for the Omanic benchmark and associated experimental implementation.

XiaojieGu/Omanic dataset
Open GitHub

Repository for the ACL 2025 OMGM coarse-to-fine multimodal retrieval and RAG framework.

ChaoLinAViy/OMGM framework
Open GitHub

Repository for the OmniCode benchmark code and data.

seal-research/OmniCode dataset
Open GitHub

Repository associated with the paper's implementation of probabilistic monolithic and model reconciling explanation algorithms.

YODA-Lab/Probabilistic-Monolithic-Model-Reconciling-Explanations benchmark
Open GitHub

Repository stated by the paper as containing code to reproduce the experiments for the uncertainty-measure framework.

ml-jku/uncertainty-measures implementation
Open GitHub

Repository linked by the paper as the code for the controlled shortest-path reasoning experiments.

riccardoalberghi/DP benchmark
Open GitHub

Public ReAct repository used as the source code-base for the original ReAct setup and corresponding experiments.

ysymyth/ReAct framework
Open GitHub

Repository containing the RepoExec benchmark/source code for executable repository-level code generation evaluation and related dependency-utilization tooling.

FSoft-AI4Code/RepoExec dataset
Open GitHub

Repository containing code for the paper, including configurations for BM25 baseline, ReAct agent variants, self-reflection, AutoCodeRover-related settings, and evaluation scripts.

JetBrains-Research/ai-agents-code-editing benchmark
Open GitHub

Repository containing the ARC task data and supporting materials for the benchmark introduced and analyzed in the paper.

fchollet/ARC dataset
Open GitHub

Code repository for the paper's SimpleLogic experiments and constructive BERT reasoning implementation.

joshuacnf/paradox-learning2reason
Open GitHub

Microsoft's TorchScale repository exposes X-MOE as a configurable sparse-MoE feature and includes this paper in its citation list.

microsoft/torchscale implementation
Open GitHub

Repository containing AutoTransform, AutoInject, Challenger/Inspector defenses, experiments, figures, and benchmark resources for the paper.

CUHK-ARISE/MAS-Resilience
Open GitHub

Repository for the paper's PyTorch implementation, including scripts, prompts, helper functions, and instructions for running memory representation and retrieval experiments.

zengrh3/StructuralMemory dataset
Open GitHub

Repository for the paper's model preparation, device-specific conversion and deployment, unified inference measurement, power parsing, and benchmark result artifacts.

PINetDalhousie/EdgeAI-Sustainability
Open GitHub

Official ToolBench repository containing benchmark tasks, action-generator and evaluator code, tests, setup instructions, and evaluation commands.

sambanova/toolbench benchmark
Open GitHub

Repository path for Agent Spec runtime adapters that translate Agent Spec components into framework-specific equivalents for popular agentic frameworks.

oracle/agent-spec implementation
Open GitHub

WayFlow is presented as the paper's reference runtime for executing Agent Spec components, including native support for Agent Spec Agents and Flows.

oracle/wayflow framework
Open GitHub

Repository announced by the authors for code and data supporting the neutral event graph induction framework.

liusiyi641/Neutral-Event-Graph dataset
Open GitHub

Open-source code and data repository for the OpenAgentSafety framework and benchmark task suite.

Open-Agent-Safety/OpenAgentSafety dataset
Open GitHub

Official implementation repository for OpenAI Gym, the reinforcement-learning environment toolkit introduced and described by the paper.

openai/gym
Open GitHub

OpenAI Evals implementation for measuring language-model text compression and decompression behavior, including the possibility of hiding information in compact strings.

openai/evals
Open GitHub

Repository linked by the authors as the location of all generated code and other answers used in the study.

lmous/openai-gpt4-coding-assistant implementation
Open GitHub

OpenAssistant's open-source repository containing the project code, data-collection application components, guidelines/configuration, and model-training implementation associated with OASST1.

LAION-AI/Open-Assistant dataset
Open GitHub

The open-source implementation of the OpenHands platform introduced and evaluated in the paper.

All-Hands-AI/OpenHands
Open GitHub

Meta AI's metaseq repository contains the codebase used to train and experiment with OPT models, including OPT-specific project materials.

facebookresearch/metaseq
Open GitHub

A modular, parallelized code base for simulating constant-product AMM trading against a CEX, designed for large-scale experiments and extensible market-design variants.

JasonSome/cpmm-trading implementation
Open GitHub

Library of ADMM applications for sparse and low-rank optimization used to test NewADMM.

canyilu/LibADMM package
Open GitHub

Huawei Cloud VM-placement traces used in the cloud resource scheduling case study.

huaweicloud/VM-placement-dataset dataset
Open GitHub

Official implementation of OCTree for optimized feature generation for tabular data via LLMs with decision-tree reasoning.

jaehyun513/OCTree framework
Open GitHub

Official repository linked by the paper for the ORACLE framework introduced and evaluated in the study.

yangzhj53/ORACLE
Open GitHub

Open-source MA-Gym codebase for evaluating Manager Agents in graph-based multi-agent workflow orchestration.

DeepFlow-research/manager_agent_gym framework
Open GitHub

Open-Finance-Lab AgenticTrading is an open-source experimental playground for LLM-powered trading agents, backtests, paper-trading simulations, reasoning logs, benchmark comparisons, and the FinAgent orchestration subsystem.

Open-Finance-Lab/AgenticTrading framework
Open GitHub

Repository for the OrderFusion model and workflow, including package installation, tutorial notebook, data-reading, model optimization, evaluation, and forecast-plotting functions.

runyao-yu/OrderFusion package
Open GitHub

Repository for Orla, the library and execution engine for constructing and serving LLM-based agentic workflows with stage mapping, orchestration, and workflow-level memory management.

dorcha-inc/orla framework
Open GitHub

Official implementation of PaCA, including a modified PEFT library and examples for task-specific fine-tuning, instruction tuning, NVIDIA GPUs, and Intel Gaudi HPUs.

WooSunghyeon/paca
Open GitHub

Official PackKV repository containing custom C++/CUDA extensions, model integrations, evaluation code, scripts, and reproducibility commands.

BoJiang03/PackKV
Open GitHub

Open-source repository for the pAI/MSc research agent and workflow described in the technical report.

PoggioAI/PoggioAI_MSc
Open GitHub

Claude Code Skill companion for PoggioAI/MSc.

PoggioAI/PoggioAI_MSc-claude
Open GitHub

Official Uber Research repository containing implementations of the original POET and Enhanced POET algorithms.

uber-research/poet
Open GitHub

Official Palu implementation with compression and Fisher rank-search code, perplexity and zero-shot evaluation, LongBench evaluation, and attention/reconstruction latency kernels.

shadowpa0327/Palu
Open GitHub

Public source-code repository for PAMS, the Python-based Platform for Artificial Market Simulations.

masanorihirano/pams framework
Open GitHub

Official code artifact for the paper, including the PEFT/S4 adapter configuration and an example command for evaluating it with a RoBERTa backbone on GLUE.

amazon-science/peft-design-spaces
Open GitHub

Repository containing BART and T5 training code, LoRA components, layer-retention configurations, dataset preprocessing scripts, and experiment configuration files for the paper.

zhuyunqi96/LoraLPrun
Open GitHub

Repository for the RedPajama data recipe and corpus from which the paper draws compute-dependent pretraining subsets.

togethercomputer/RedPajama-Data dataset
Open GitHub

Author repository containing code and instructions to reproduce the BERT passage re-ranking experiments on MS MARCO and TREC-CAR.

nyu-dl/dl4marco-bert
Open GitHub

Open-source implementation of the 2D odd-one-out environments whose basic, confounded, and experimenting/meta-learning variants are used and adapted in the paper's passive-agent experiments.

deepmind/tell_me_why_explanations_rl
Open GitHub

PathBench contains shared implementations and configurations for WSI-Caption, HistGen, BiGen, and SCOUT, fixed report splits, conventional NLG metrics, and dataset-specific CRQS pipelines.

surykntsingh/PathBench benchmark
Open GitHub

Repository provided by the authors to support reproduction of the Identifier-Organizer-Adapter data synthesis and distillation framework.

BokwaiHo/IOA framework
Open GitHub

Official open-source repository for the PedNStream pedestrian-network simulator, including core LTM modules, scenario data, examples, visualization, tests, and configuration files.

WaimenMak/PedNStream
Open GitHub

Official ParlAI repository containing the Persona-Chat task infrastructure and associated dialogue-model resources.

facebookresearch/ParlAI
Open GitHub

Code used to scrape and process the dataset, construct train/test and benchmark datasets, and build and evaluate baseline phishing detection models.

phreshphish/phreshphish dataset
Open GitHub

PKU-Alignment's Safe-RLHF/Beaver repository implements SFT, reward/cost modeling, RLHF, and SafeRLHF and explicitly announces/releases the PKU-SafeRLHF dataset family through its project resources.

PKU-Alignment/safe-rlhf
Open GitHub

Repository linked by the paper for the updated PlanBench benchmark, including tools, datasets, prompt/result reproduction scripts, planning utilities, and the PlanBench implementation.

karthikv792/LLMs-Planning dataset
Open GitHub

CMBAgent Benchmarks repository used for CAMB tool-grounded precision tasks.

cmbagent/Benchmarks dataset
Open GitHub

Repository containing the implementation of the Point-M2AE hierarchical point-cloud pre-training framework.

ZrrSkywalker/Point-M2AE implementation
Open GitHub

Official codebase for Policy Decorator, including offline base-policy training and online residual-refinement scripts for ManiSkill and Adroit configurations.

tongzhoumu/policy_decorator
Open GitHub

Agent Network Protocol is treated as an open-network agent discovery and collaboration protocol using decentralized identifiers and JSON-LD.

agent-network-protocol/AgentNetworkProtocol
Open GitHub

Model Context Protocol is treated as a standardized context-ingestion and tool-invocation protocol relevant to execution-level transitions.

modelcontextprotocol
Open GitHub

Official repository released by the authors for the paper's post-training calibration experiments and PosConf method.

EIT-NLP/Post-Training-Calibration
Open GitHub

Contains framework code, domain profiles, prompt builders, placeholder QA, deterministic replacement, paired full-resolution benchmark posters, evaluation outputs, runtime and failure audits, prompts, manifests, and configuration records.

tyy99phy/paper_poster_harness
Open GitHub

Repository associated with the paper's dataset and experimental framework for evaluating LLM-generated scientific reviews against human reviews and post-publication outcomes.

akhilpandey95/LMRSD dataset
Open GitHub

Repository containing code for the paper, including PhraseBank preprocessing, training, and testing scripts for the LLaMA financial sentiment analysis experiments.

luosting/LLaMA-Financial-sentiment-analysis implementation
Open GitHub

Open-source code for the Hypothetical Document Embeddings retrieval method introduced and evaluated in the paper.

texttron/hyde
Open GitHub

Apache Airflow repository; used as an example among ten large open-source industrial projects from which function-summary pairs were sampled.

apache/airflow
Open GitHub

RxJava repository; used as an example among ten large open-source industrial projects from which function-summary pairs were sampled.

ReactiveX/RxJava
Open GitHub

Official repository containing the raw Provo data used by the study, generation scripts and saved model continuations, uncertainty estimates, joint analysis notebooks, semantic and syntactic TVD experiments, and appendix experiments.

evgeniael/predict_next_word
Open GitHub

Repository released by the authors for Step Law experimental code, data/configuration artifacts, and model checkpoints.

step-law/steplaw
Open GitHub

Original StockNet codebase for stock movement prediction from tweets and historical prices.

yumoxu/stocknet-code benchmark
Open GitHub

Repository for the StockNet dataset, containing historical stock-price data and tweet-data structure for stock movement prediction from tweets and historical prices.

yumoxu/stocknet-dataset dataset
Open GitHub

Repository reported by the paper as the source code for the LLM-enhanced tweet emotion analysis and stock movement prediction framework.

anv0101/stock-prediction implementation
Open GitHub

LLT R package used by the authors to transform the cryptocurrency datasets before classifier evaluation.

mtkurbucz/LLT package
Open GitHub

A Python benchmark for backtesting prediction-market trading agents using real Kalshi market replay data, included episodes, agent interfaces, simulator configuration, and metrics output methods.

Oddpool/PredictionMarketBench dataset
Open GitHub

Official code repository for the prefix-tuning method and experiments introduced by the paper.

XiangLi1999/PrefixTuning
Open GitHub

Repository containing the generated historical_traffic_data.csv dataset, location.py simulation code, and HTML traffic-map and prediction visualizations used by the paper.

Psxxg/Fairly-Private
Open GitHub

Repository containing the ProAgent code, environment setup, experiment scripts, prompt/agent implementation, and instructions for reproducing the Overcooked-AI evaluations.

PKU-Alignment/ProAgent
Open GitHub

Official implementation repository for ProAgent, the LLM-based agent system introduced by the paper for Agentic Process Automation.

OpenBMB/ProAgent framework
Open GitHub

Official codebase containing prompt templates, model/evaluation utilities, scripts to reproduce figures and significance tests, a NiceGUI corpus visualization app, and download links for the OEDD v1.0.0 test corpus and initial results.

sonnygeorge/OEDD dataset
Open GitHub

Public code repository for the trajectory-probing experiments and analyses introduced in the paper.

AndresAlgaba/probing_reasoning_traces benchmark
Open GitHub

Repository containing task templates, benchmark-generation scripts, preprocessing, model-prediction configurations, response extraction, and evaluation code for reproducing and extending the ProcBench experiments.

ifujisawa/proc-bench benchmark
Open GitHub

Repository containing code, prompts, and data for reproducing or evaluating Program of Thoughts prompting.

wenhuchen/Program-of-Thoughts benchmark
Open GitHub

Official repository containing generated self-play data, fine-tuning datasets, evaluation directories, notebooks, and supporting code for the SInQ experiments.

Avmb/semantic_neq_game
Open GitHub

Official repository location for the MBPP programming-problem benchmark introduced by the paper.

google-research/google-research dataset
Open GitHub

Notebook implementing the MathQA-to-Python translation and generation workflow used for MathQA-Python.

google/trax implementation
Open GitHub

Official ProgramBench repository for the benchmark, tooling, usage guide, and baseline evaluation workflow.

facebookresearch/ProgramBench dataset
Open GitHub

Repository containing the progressive multimodal search-agent rollout, image and text tools, retrieval and summarization services, TN-GSPO modifications, training scripts, evaluation scripts, and dataset configuration files.

DingWu1021/Promsa framework
Open GitHub

Official ProtAgents repository containing the agent implementation, shared functions, ProteinForceGPT integration code, configuration utilities, and notebooks corresponding to Experiments I, II, and III.

lamm-mit/ProtAgents
Open GitHub

Repository for the paper's trace-rewriting experiments, including configuration files, source scripts for trace generation, rewriting, distillation and evaluation, prompt-optimization scripts, and pre-generated GSM8K/MATH rewritten trace datasets.

xhOwenMa/trace-rewriting dataset
Open GitHub

Repository for SimpleVLA-RL, a reinforcement-based VLA approach that uses online learning with a binary reward and a single demonstration trajectory.

PRIME-RL/SimpleVLA-RL
Open GitHub

Repository released by the authors for the PuzzleVQA benchmark, including the research artifact supporting dataset generation and evaluation.

declare-lab/LLM-PuzzleTest dataset
Open GitHub

Repository for evaluating LLM agents on real-world coding tasks, with benchmark code, configuration, data folders, unit-test workflow, PyInstruct data link, and PyLlama3 model link.

Mercury7353/PyBench dataset
Open GitHub

Repository containing the public implementation and data-processing pipeline associated with the paper.

DARE-ML/DeepGR4J-Extremes framework
Open GitHub

Repository containing Qlib's code, documentation, and additional platform features beyond those described in the paper.

microsoft/qlib framework
Open GitHub

Official QLoRA implementation and training code released by the paper authors.

artidoro/qlora implementation
Open GitHub

bitsandbytes library containing low-bit quantization and CUDA components used to implement QLoRA's 4-bit training stack.

TimDettmers/bitsandbytes
Open GitHub

Repository maintaining QMSum data, train/validation/test splits, domain-specific data folders, extracted spans, model outputs, figures, and a data-processing notebook.

Yale-LILY/QMSum dataset
Open GitHub

Official implementation and reproducibility artifact for QSpec, including installation instructions, Docker support, compiled kernels, experiment scripts, and vLLM-related code.

hku-netexplo-lab/QSpec
Open GitHub

Python and shell-script repository containing QTMRL or multi-indicator experiment code, non-multi-indicator model scripts, baseline scripts, dependency specifications, and instructions for downloading a related multi-indicator dataset.

ChenJiahaoJNU/QTMRL framework
Open GitHub

Code and example workflow for EcoDatum's operator outputs, labeling functions, LabelModel ensemble, configuration, and curated JSONL generation.

haomo-ai/EcoDatum
Open GitHub

Repository containing paper versions, an overview image, README, and Jupyter Notebook implementation for Quantformer.

zhangmordred/QuantFormer implementation
Open GitHub

Repository cited for GPT-J-6B, the 6B autoregressive language model evaluated as the largest model in the paper's main Pile-based scaling experiments.

kingoflolz/mesh-transformer-jax
Open GitHub

Repository containing code and data supporting the RAG application, custom probes and detectors, Garak configuration, synthetic PII data, and thesis experiments.

rhmoult/DSU
Open GitHub

Repository containing code, data, model implementations, visualisations, and setup instructions for classic and quantile versions of linear regression, BD-LSTM, Conv-LSTM, and ED-LSTM across BTC, ETH, Sunspot, Mackey-Glass, and Lorenz datasets.

sydney-machine-learning/quantiledeeplearning dataset
Open GitHub

Official Quark implementation with separate branches for toxicity unlearning, sentiment steering, and repetition reduction, plus evaluation scripts and released checkpoint links.

GXimingLu/Quark
Open GitHub

Contains the contextual chunker, embedding and reranker training code, Kaggle submission notebooks, official competition data layout, and experiments corresponding to the baseline and final system.

nuinashco/unlp2026_shared_task system
Open GitHub

Official repository for the Qwen-VLA model and technical report, containing project information, demonstrations, and reported benchmark summaries.

QwenLM/Qwen-VLA
Open GitHub

Official repository for the Qwen3 Embedding and Qwen3 Reranker research artifacts introduced in the paper.

QwenLM/Qwen3-Embedding
Open GitHub

Official Qwen3 repository associated with the model family introduced in the report.

QwenLM/Qwen3
Open GitHub

LMFlow contains the paper's RAFT alignment implementation, including a run_raft_align.sh example and diffusion demonstrations.

OptimalScale/LMFlow
Open GitHub

Adrian H. Raudaschl's RAG-Fusion repository, cited by the paper as the implementation/artifact context for the RAG-Fusion workflow.

Raudaschl/RAG-Fusion
Open GitHub

Repository containing the RAGSmith codebase and the datasets used in the study.

yAquila/RAGSmith dataset
Open GitHub

Repository released by the authors containing the SGLang-based implementation of RLT and LBGR together with scripts and configurations for reproducing the GSP, ShareGPT, UltraChat, and Loogle evaluations.

fzwark/KVRouting
Open GitHub

Author-maintained repository containing Python and MATLAB implementations and examples for randomizing affine-diffusion models and computing randomized characteristic-function constructions.

LechGrzelak/Randomization implementation
Open GitHub

Repository associated with Rational Tuning experiments for LLM cascade modeling and threshold optimization.

mzelling/rational-llm-cascades implementation
Open GitHub

Official repository containing the RDMA implementation and the datasets generated or analyzed for the study.

jhnwu3/RDMA
Open GitHub

Official RE-Bench repository containing the seven task families and supporting setup material used to evaluate autonomous AI R&D capabilities.

METR/RE-Bench benchmark
Open GitHub

Public implementation of the Modular agent scaffold used as one of the two main agent designs in RE-Bench experiments.

poking-agents/modular-public
Open GitHub

Open-source AIDE machine-learning engineering agent, lightly adapted as the second main scaffold evaluated in RE-Bench.

WecoAI/aideml
Open GitHub

Alpaca Eval is one of the two central input sets compared in the paper's automatic bencher experiments.

tatsu-lab/alpaca_eval dataset
Open GitHub

Repository for the paper's released code and data supporting the RealRank automatic LLM ranking analysis.

yale-nlp/RealRank benchmark
Open GitHub

CHEF is one of the three real-world datasets used to evaluate STEEL.

THU-BPM/CHEF dataset
Open GitHub

CrewAI is presented as a Python framework for defining LLM-based agents, tasks, tools, sequential execution, entity memory, and callbacks.

crewAIInc/crewAI framework
Open GitHub

Jupyter notebook implementing a workflow that performs root-cause analysis from a directly-follows graph abstraction and adds evaluation steps such as confidence scoring and reasoning output.

fit-alessandro-berti/agents-trial implementation
Open GitHub

Jupyter notebook implementing the process-mining fairness workflow with protected-group identification and comparison between protected and non-protected cases.

fit-alessandro-berti/agents-trial implementation
Open GitHub

Official Chronos repository for pretrained time-series forecasting models.

amazon-science/chronos-forecasting implementation
Open GitHub

Datadog Toto repository for Time-Series-Optimized Transformer for Observability.

DataDog/toto implementation
Open GitHub

Official Google Research TimesFM repository for the Time Series Foundation Model.

google-research/timesfm implementation
Open GitHub

IBM Granite TSFM TinyTimeMixer model implementation.

ibm-granite/granite-tsfm implementation
Open GitHub

MOMENT repository for a family of open time-series foundation models.

moment-timeseries-foundation-model/moment implementation
Open GitHub

TiRex repository for zero-shot forecasting across long and short horizons.

NX-AI/tirex implementation
Open GitHub

Moirai/Uni2TS repository for universal time-series forecasting Transformers.

SalesforceAIResearch/uni2ts implementation
Open GitHub

Sundial repository for highly capable time-series foundation models.

thuml/Sundial implementation
Open GitHub

Lag-Llama repository for probabilistic time-series foundation forecasting.

time-series-foundation-models/lag-llama implementation
Open GitHub

Repository for the Re4 Scientific Computing Agent, including source code in Jupyter Notebook and Python script formats and a README describing the rewriting-resolution-review-revision logical chain.

ChengAo21/Re4_Sci_Agent framework
Open GitHub

Repository for the ICLR 2023 ReAct prompting paper, including data, prompts, HotpotQA, FEVER, ALFWorld, and WebShop notebooks, plus Wikipedia environment wrappers.

ysymyth/ReAct benchmark
Open GitHub

Repository for the paper's training-data synthesis implementation.

BAAI-DCAI/Training-Data-Synthesis implementation
Open GitHub

GitHub repository associated with the REAL websites, framework, and leaderboard for benchmarking autonomous web agents.

agi-inc/agisdk framework
Open GitHub

Repository containing the paper's complete demonstration code and challenge materials and referenced as the Sim-to-Real repository.

AMD-AIM/Physical_AI_Challenge
Open GitHub

Workshop implementation associated with the synthetic-data-generation and Sim-to-Real manipulation workflow.

PhysicalAI-AIM/Robot_synthetic_data_generation_workshop
Open GitHub

Open-source real-time Vulkan 3D Gaussian Splatting renderer/plugin used to implement the reconstruction-to-simulation bridge described by the Real2Sim pipeline.

ZJLi2013/vk-gsplat-plugin
Open GitHub

Google Research code for REALM pre-training, document-index refreshing, example generation, released checkpoints, and integration with the ORQA fine-tuning code.

google-research/language implementation
Open GitHub

Repository containing the yield-farming simulation environment used for strategy analysis.

xujiahuayz/yieldAggregators implementation
Open GitHub

Official code repository for reproducing the paper's Game of 24, ALFWorld, BlocksWorld, and Tic-Tac-Toe experiments.

agentification/RAFA_code
Open GitHub

Official implementation and release repository for the DramaSR-LRM pipeline, benchmark data structure, training, inference, evaluation, and model checkpoints.

198808xc/DramaSR-LRM benchmark
Open GitHub

GitHub directory containing REVEAL code, AIGC-text-bank files, configurations, inference scripts, training scripts, and setup instructions.

microsoft/AnthropomorphicIntelligence dataset
Open GitHub

Official repository maintained by the authors to collect and update literature and resources on static and dynamic LLM benchmarking under data contamination.

SeekingDream/Static-to-Dynamic-LLMEval
Open GitHub

The ParlAI repository contains the framework, tasks, model access, fine-tuning and evaluation code used to reproduce and extend the paper's chatbot recipes.

facebookresearch/ParlAI
Open GitHub

Repository released by the authors for ReCode, including the framework and associated research artifacts.

ZJU-CTAG/ReCode dataset
Open GitHub

Microsoft's RecAI research repository contains the InteRecAgent implementation in a dedicated subdirectory and describes the paper's LLM-as-brain/recommender-models-as-tools system.

microsoft/RecAI
Open GitHub

Official repository accompanying the paper, with scripts for threshold sensitivity, context, typo, adversarial-prompt, and long-form experiments.

duygunuryldz/uncertainty_in_the_wild
Open GitHub

GitHub repository for the benchmark code and evaluation tooling introduced by the paper.

JiseungHong/ReCUBE benchmark
Open GitHub

Repository released by the authors for LLM-Rsum, with code paths for ChatGPT/Llama experiments, recursive memory, retrieval, MemoChat, and MemoryBank comparisons.

qingyue2014/Rsum
Open GitHub

Official repository for the paper, containing Red-Eval prompt/evaluation code, harmful-question resources, result files, and Starling safety-alignment training code.

declare-lab/red-instruct
Open GitHub

Code and data repository for the RedAgent framework, including context probing, memory, adaptive routing, evaluation, and example workflows.

QuirkShark/RedAgent
Open GitHub

Repository for the RedKnot serving runtime, including the head-aware KV-reuse, Elastic Sparsity, SegPagedAttention, model-adapter, and benchmark implementation.

rednote-machine-learning/RedKnot
Open GitHub

Repository containing code for preparing the RedPajama datasets and reproducing the data-processing workflow described by the paper.

togethercomputer/RedPajama-Data dataset
Open GitHub

HuggingFace Diffusers repository, used for refactoring diffusion-model implementation files such as UNet and scheduler sources.

huggingface/diffusers
Open GitHub

Official repository released by the authors with implementations, demos, prompts, and research artifacts for Reflexion experiments.

noahshinn024/reflexion
Open GitHub

Repository containing the implementation associated with the regime-aware continual adaptive portfolio-management framework.

Dumail/ReCAP framework
Open GitHub

Repository associated with the paper's Qwen2.5-0.5B SFT, DPO, RLOO, reward-model, Countdown, and external-verifier experiments.

Yifu93/LLM-Reinforcement-Learning implementation
Open GitHub

Code repository for the Relational Representation Distillation implementation reported by the paper.

giakoumoglou/distillers implementation
Open GitHub

Public source code for the paper's RPS-based ordinal conformal prediction method and experiments.

stefanahaas41/rps-ordinal-conformal-prediction
Open GitHub

Repository maintaining extended notes, open research items, and per-claim literature verdicts for the removable-defects project.

macrokit/removable-defects
Open GitHub

The Arena Hard repository is used as the pairwise chatbot evaluation setup; the authors generated additional model outputs and scored them with the compared judges.

lm-sys/arena-hard benchmark
Open GitHub

The public implementation of RepoAgent, an LLM-powered repository agent for generating, maintaining, updating, and understanding repository-level documentation.

OpenBMB/RepoAgent framework
Open GitHub

Official RepoBench repository containing data, evaluation scripts, experiment code, and documentation for the benchmark introduced by the paper.

Leolty/repobench dataset
Open GitHub

Repository containing RepoGraph code for constructing and retrieving repository graph context, plus integrations with Agentless and SWE-agent and scripts for SWE-bench evaluation.

ozyyshr/RepoGraph framework
Open GitHub

RepoReviewer repository containing the Python backend, Next.js frontend, screenshots, CLI/API/web workflow, and templates or utilities for future empirical evaluation and annotation.

peng1z/RepoReviewer framework
Open GitHub

Official codebase for training ReProbe/UHead-style verifiers, generating annotated reasoning datasets, evaluating ReProbe and baselines, and reproducing benchmark tables.

ReProbe/ReProbe benchmark
Open GitHub

GitHub repository containing data from the paper's experiments, provided to support reproducibility of the proposed evaluation approach.

andstor/agentic-ai-eval-replication-package implementation
Open GitHub

Code repository for the paper's deep-search reranking and Effective Token Cost experiments.

sahel-sh/DeepHone
Open GitHub

Official code release with command-line tools, prompts, metrics, validation utilities, tests, and a small smoke-test split for the three ResearchBench tasks.

ankitala/ResearchBench benchmark
Open GitHub

Official repository for the paper, containing throughput and length experiments, negative-sample analysis, predictor tools, serving integrations, environment files, and cached results.

LLMkvsys/rethink-kv-compression
Open GitHub

Repository containing PopQA training and test data, preprocessing artifacts, generator and reranker LoRA adapters, training and testing scripts, prompt-generation code, and utilities.

CoderrrSong/CoRAG dataset
Open GitHub

Repository for Byzantine Fault Tolerance in LLM-Based Multi-Agent Systems, including pilot experiments, prompt-level confidence probing, hidden-level confidence probing, datasets, pretrained confidence probes, and experiment scripts.

Z1ivan/Byzantine-Fault-Tolerance-in-LLM-MAS framework
Open GitHub

Official repository for the BAR-RAG training, filtering, inference, and evaluation pipeline introduced in the paper.

GasolSun36/BAR-RAG
Open GitHub

The KILT repository provides the standardized Wikipedia knowledge source used for retrieval in both dialogue benchmarks.

facebookresearch/KILT
Open GitHub

ParlAI is the dialogue research framework in which the paper trains and evaluates all models and through which the RAG-based implementations and pretrained models were released.

facebookresearch/ParlAI
Open GitHub

A curated, continuously updated collection of retrieval-augmented generation work for computer vision, organized across visual understanding, generation, documents, video, multimodal tasks, and embodied AI.

zhengxuJosh/Awesome-RAG-Vision
Open GitHub

Repository identified by the paper as containing the code used for the Polymer Literature Scholar and its RAG workflows.

Ramprasad-Group/RAG
Open GitHub

The paper's RAG implementation and experiment scripts were ported to and open-sourced within the Hugging Face Transformers repository.

huggingface/transformers
Open GitHub

Official companion repository for the survey, containing survey materials and an evolving RAG knowledge base covering papers, datasets, benchmarks, evaluation resources, and toolkits.

Tongji-KGLLM/RAG-Survey
Open GitHub

Official codebase for RetroMAE pre-training and downstream dense-retrieval fine-tuning, with pretrained model resources referenced by the paper.

staoxiao/RetroMAE
Open GitHub

Open-Finance-Lab repository for FinRL Contest 2024, linked directly in the paper alongside the contest website.

Open-Finance-Lab/FinRL_Contest_2024 framework
Open GitHub

Official repository released by the authors with code, data, and agent-generated traces for reproducing the study.

ChuanMeng/text-ranking-in-deep-research benchmark
Open GitHub

Implementation repository used for the Rank1 reasoning-based re-ranker evaluated in the pipeline and query-mismatch experiments.

orionw/rank1
Open GitHub

Repository for the BrowseComp-Plus fixed-corpus deep-research benchmark and its evaluation tooling.

texttron/BrowseComp-Plus dataset
Open GitHub

Official Qwen3 repository for the model family used in the paper's most controlled comparisons of scale, post-training, reasoning mode, dense versus MoE architecture, and quantization.

QwenLM/Qwen3
Open GitHub

Repository containing the prompts and RL training code for the Ultimatum Game, matrix games, and DealOrNoDeal domains studied in the paper.

minaek/reward_design_with_llms implementation
Open GitHub

Repository for Stackelberg Reward Shaping (SRS) and its integration with inference-time alignment methods.

Haichuan23/Stackelberg-Reward-Shaping implementation
Open GitHub

Repository containing code and experiment folders for the rewarded-soups method, including LLaMA RLHF and image-captioning reproductions.

alexrame/rewardedsoups
Open GitHub

Repository for reproducing or using RFEval, the paper's benchmark for auditing reasoning faithfulness under counterfactual reasoning intervention.

AIDASLab/RFEval dataset
Open GitHub

Repository containing the implementation and experiments for the risk-aware GUMDP framework and ERM-MCTS evaluation introduced in the paper.

gh0stwin/risk-aware-gumdp
Open GitHub

GitHub repository for the simulator used to quantify uncertainty propagation in AI-augmented systems.

EMezzi/AI-Augmented implementation
Open GitHub

Official Risky-Bench repository containing benchmark datasets, data generation scripts, evaluator components, and evaluation workflows for the paper.

SophieZheng998/Risky-Bench dataset
Open GitHub

Repository released by the paper authors for RLAIF-V.

RLHF-V/RLAIF-V
Open GitHub

Repository linked by the paper as the code artifact for Variational Alignment with Re-weighting; the repository page was empty when checked.

DuYooho/VAR implementation
Open GitHub

Repository for loading and manipulating RoboNet and for training the supervised inverse and video-prediction models used with the dataset.

SudeepDasari/RoboNet
Open GitHub

Python implementation of C-EDL and the comparative experimental framework, including supported datasets, evidential models, abstention metrics, and adversarial attacks.

team-daniel/cedl
Open GitHub

Repository containing prompts and templates used to instantiate LLM-Modulo for Travel Planner and Natural Plan domains.

Atharva-Gundawar/LLM-Modulo-prompts framework
Open GitHub

Repository containing the notebook and reproduction-oriented materials for the clean, noisy, and denoised taxi-fare comparison.

padmavathi026/Smart-Fare-Prediction implementation
Open GitHub

A project that crawls WeChat users' profile pictures with usernames and visualizes them; used as the source project for the bug-fixing task.

yangxuanxc/wechat_friends
Open GitHub

Official ROSClaw repository implementing the physical-agent runtime, robot and tool integration, safety and simulation components, execution trace capture, resource management, and supporting documentation.

ros-claw/rosclaw
Open GitHub

Repository for serving, training, and evaluating LLM routers, including router types corresponding to the paper such as matrix factorization, similarity-weighted ranking, BERT, causal LLM, and random routing.

lm-sys/routellm framework
Open GitHub

Repository released with the paper for adaptation, router-sensitivity scoring, structural pruning, and evaluation.

ianKa1/MoE_pruning
Open GitHub

Official RouterBench code repository containing data converters, prompt embeddings, predictive and cascading routers, evaluation code, configurations, tests, and visualization utilities.

withmartian/routerbench benchmark
Open GitHub

Repository for the paper's RUArt implementation, including model code, configuration, inference entry points, dependency instructions, and links to pretrained and preprocessed files.

xiaojino/RUArt
Open GitHub

Repository reported by the paper for the AtomicTranslation code used in the language-to-logic translation experiments.

KrisAesoey/AtomicTranslation dataset
Open GitHub

Repository containing the S-LoRA serving implementation, benchmark scripts, tests, and documentation for Unified Paging, heterogeneous LoRA batching, and the evaluated baselines.

S-LoRA/S-LoRA
Open GitHub

GitHub repository for a convolutional neural stock-market technical analyser used as the starting point for the paper's proposed model.

philipxjm/Convolutional-Neural-Stock-Market-Technical-Analyser implementation
Open GitHub

Repository for SafeArena, a benchmark for assessing harmful capabilities and safety risks of autonomous web agents.

McGill-NLP/safearena dataset
Open GitHub

Repository containing code and prompts used when writing the paper, including agent setup notebooks, safety architecture notebooks, image-generation safety notebooks, requirements, utilities, and unsafe agent test requests.

ishaandomkundwar/Agent-Safety dataset
Open GitHub

Repository for the SafeRAG benchmark, including attack-task data/knowledge bases, retrievers, prompts, metrics, evaluators, model interfaces, and quick-start evaluation code.

IAAR-Shanghai/SafeRAG dataset
Open GitHub

Google Gemini full-stack LangGraph quickstart prototype used as the basis for a query-decomposition, search, reflection, and answer-finalisation scaffold.

google-gemini/gemini-fullstack-langgraph-quickstart system
Open GitHub

Repository for SafeSearch code, dataset, prompts, assets, and red-teaming configurations.

jianshuod/SafeSearch dataset
Open GitHub

Official implementation of the Vera safety-testing framework, including taxonomy exploration, case generation, adaptive execution, Vera-Bench, and guard-model fine-tuning components.

Yunhao-Feng/Vera framework
Open GitHub

DeepSpeed repository containing the distributed training framework and MoE functionality introduced and evaluated by the paper.

microsoft/DeepSpeed framework
Open GitHub

Repository containing PyTorch DiT model definitions, pretrained ImageNet checkpoints, and training and sampling code corresponding to the paper.

facebookresearch/DiT
Open GitHub

Revised Tau2-Bench repository used by the authors because they describe the original benchmark as containing noisy data.

AGI-Eval-Official/tau2-bench-revised benchmark
Open GitHub

Public Flan-T5 checkpoints released with the paper in the T5X model repository.

google-research/t5x
Open GitHub

Repository containing the MMLU chain-of-thought prompts used for the paper's MMLU evaluation.

jasonwei20/flan-2
Open GitHub

Repository for BIG-bench, from which the paper evaluates 62 tasks covering reasoning, knowledge, social behaviour, and other language-model capabilities.

google/BIG-bench benchmark
Open GitHub

LLM-Random research codebase containing model-training infrastructure and research configurations associated with the fine-grained MoE scaling-law experiments.

llm-random/llm-random implementation
Open GitHub

BigSparse project code associated with the paper's sparsity experiments and scaling-law fitting workflow.

google-research/jaxpruner
Open GitHub

NL2Flow is the automated workflow problem generation and symbolic evaluation framework used to generate planning problems, compile PDDL, and evaluate plans.

IBM/nl2flow framework
Open GitHub

NL2FLOW-Runner contains code for running the experiments reported in the paper.

IBM/nl2flow-runner implementation
Open GitHub

Repository containing BiomedGPT training, fine-tuning, preprocessing, dataset documentation, and model-use resources associated with the original and extended BiomedGPT research line.

taokz/BiomedGPT implementation
Open GitHub

Repository containing canonical tool contracts, interface condition renderers, validation and structured diagnostics, deterministic sandbox executors, run matrix harnesses, structured logging, and aggregate metric scripts.

akgitrepos/schema-first-tool-apis-experiments benchmark
Open GitHub

OpenHands CodeAct is used as one of the agent frameworks evaluated on ScienceAgentBench, alongside direct prompting and self-debug.

All-Hands-AI/OpenHands framework
Open GitHub

Repository containing code and data access instructions for ScienceAgentBench, including benchmark structure, agent/evaluation scripts, and links to benchmark materials.

OSU-NLP-Group/ScienceAgentBench dataset
Open GitHub

Repository for the Calo-VQ model used by SciFi to reproduce a calorimeter simulation inference and plotting pipeline.

qibin2020/calo-VQ implementation
Open GitHub

Repository containing implementation details for the SciFi autonomous agentic scientific workflow framework.

qibin2020/scifi framework
Open GitHub

Repository containing SciPredict data, configurations, evaluation scripts, background-knowledge pipelines, MCQ-to-free-form conversion, and experiment-running instructions.

scaleapi/scipredict dataset
Open GitHub

Official repository containing the SCIZOR curation code, customized Octo and RoboMimic training paths, and scripts for progress-based scoring and deduplication.

UT-Austin-RPL/SCIZOR
Open GitHub

Hosts training and inference scripts, SeaBench-Audio materials, project assets, and links to the released model and benchmark.

DAMO-NLP-SG/SeaLLMs-Audio benchmark
Open GitHub

Open-source Search-R1 training and inference code for reasoning-and-search interleaved LLMs, including PPO/GRPO training and retrieval integration.

PeterGriffinJin/Search-R1 implementation
Open GitHub

Repository for Reasoning-Reinforced Representation for Search, with release-status information and a linked model artifact.

ytgui/Search-R3 framework
Open GitHub

Public repository containing the SecRepoBench implementation, benchmark metadata, descriptions, harnesses, tools, and scripts for running inference and evaluation.

ai-sec-lab/SecRepoBench dataset
Open GitHub

Repository released by the authors for RLHF technical artifacts, reward models, and PPO/PPO-max code.

OpenLMLab/MOSS-RLHF implementation
Open GitHub

C-Eval benchmark repository used to assess whether PPO alignment degrades Chinese language-understanding capabilities.

SJTU-LIT/ceval benchmark
Open GitHub

Provides training code, HH-RLHF data with preference-strength information, and the GPT-4-cleaned validation set.

OpenLMLab/MOSS-RLHF dataset
Open GitHub

Official code repository containing the implementation needed to reproduce the GazeReward framework and experiments.

Telefonica-Scientific-Research/gaze_reward framework
Open GitHub

Public implementation repository for Pythia, the modular hallucination-monitoring system evaluated and discussed in the paper.

wisecubeai/pythia
Open GitHub

Repository containing code for data generation and risk-control experiments associated with selective conformal classification.

git4review/conformal_selective_classification implementation
Open GitHub

Repository linked from the paper's official project page; it contains MATRIX simulation source code, an example runner, requirements, and released MATRIX-generated alignment data.

ShuoTang123/MATRIX implementation
Open GitHub

Official repository for the paper's self-collaboration code-generation framework, containing generation/evaluation scripts, core code, data, and evaluation utilities.

YihongDong/Self-collaboration-Code-Generation implementation
Open GitHub

Repository location for the open-source UL2 model checkpoints used as one of the evaluated language models.

google-research/google-research
Open GitHub

BIG-bench StrategyQA task repository used for the paper's question-only StrategyQA evaluation set.

google/BIG-bench benchmark
Open GitHub

Contains data-processing, SFT, iterative reward-model training and filtering, inference, GPT-4 evaluation, and PPO code for Self-Evolved Reward Learning.

microsoft/DKI_LLM framework
Open GitHub

Repository for Self-Evolving GPT; the README states that se_gpt_MAIN.py shows the main workflow of the framework.

ArrogantL/se_gpt framework
Open GitHub

Official code and data repository for generating Self-Instruct data, classifying tasks, generating instances, filtering and formatting data, fine-tuning GPT3, and evaluating on the released user-oriented tasks.

yizhongw/self-instruct dataset
Open GitHub

Repository dedicated to the paper's SKR experiments, with data, chain-of-thought results, and scripts for SKR_prompt, SKR_icl, SKR_cls, and SKR_knn.

THUNLP-MT/SKR
Open GitHub

Official code repository for Self-Play Preference Optimization, including training, generation, ranking, and evaluation scripts and released model references.

uclaml/SPPO
Open GitHub

Repository containing the original Self-RAG implementation, critic and generator data-creation workflows, retriever setup, training scripts, and short- and long-form evaluation code.

AkariAsai/self-rag
Open GitHub

Official repository for the SELF-REFINE framework, containing code, prompts, data, task runners, evaluation scripts, and examples for the paper's studied tasks.

madaan/self-refine
Open GitHub

Repository for AlpacaEval, the automatic instruction-following evaluation framework used for the paper's AlpacaEval 2.0 leaderboard evaluation.

tatsu-lab/alpaca_eval benchmark
Open GitHub

Official PyTorch codebase for the Image-based Joint-Embedding Predictive Architecture, including training code, masks, model definitions, configurations, checkpoints, and logs.

facebookresearch/ijepa
Open GitHub

Repository for the SelfAI system, including its experiment-management software and implementation-oriented documentation.

XiaoXiao-Woo/SelfAI
Open GitHub

Contains the paper's hybrid-retrieval defense, detection implementations, evaluation and experiment scripts, result files, sanitized examples, corpus setup instructions, figures, and reproducibility documentation.

scthornton/semantic-chameleon system
Open GitHub

Contains the pipeline for loading CoQA and TriviaQA, generating answers, clustering semantic similarities, computing likelihoods and uncertainty measures, evaluating AUROC, and reproducing the paper's analyses, together with the released hand-labelled semantic-equivalence data.

lorenzkuhn/semantic_uncertainty dataset
Open GitHub

TensorFlow lm_1b repository cited for the released CNN-BIG-LSTM model and training recipes from Jozefowicz et al. (2016).

tensorflow/models implementation
Open GitHub

Code repository linked by the paper for Sentence-BERT and sentence-transformer models.

UKPLab/sentence-transformers
Open GitHub

The paper's released Sequence Tutor implementation within TensorFlow Magenta, including a checkpointed melody RNN.

tensorflow/magenta
Open GitHub

Open-source Azure production trace repository used as an evaluation data source.

Azure/AzurePublicDataset dataset
Open GitHub

The repository implements an SGDE API and client for training, uploading, downloading, and using generative models, and includes a demonstration notebook and model examples.

archettialberto/SGDE
Open GitHub

Code repository for the ShapG method introduced in the paper.

vectorsss/shapG implementation
Open GitHub

Self-service snack bar kiosk system repository containing requirements documentation, automated Robot Framework acceptance tests, C4 architecture documentation, ADRs, and traceability audit material.

JYU-GENIUS-project/snackbar_v1 implementation
Open GitHub

First iteration for vibe coding in the study; implements a snack kiosk system using React, TypeScript, PostgreSQL, Node, and Express.

JYU-GENIUS-project/VibeCode1 implementation
Open GitHub

Lovable-generated campus-treats prototype used as the unstructured vibe coding example.

Lovablekokeilu/campus-treats implementation
Open GitHub

Official codebase implementing the paper's Shifting Attention to Relevance uncertainty-estimation pipeline and dataset/model scripts.

jinhaoduan/SAR implementation
Open GitHub

Official PyTorch implementation of the Sim2Seg visual translator and RL training pipeline, with configurations, models, sample environments, and training instructions.

rll-research/sim2seg
Open GitHub

Repository for SimCSE code and pretrained sentence-embedding models.

princeton-nlp/SimCSE implementation
Open GitHub

Official Microsoft implementation of SimMIM with training and fine-tuning code, configurations, and released pretrained/fine-tuned model resources.

microsoft/SimMIM implementation
Open GitHub

Open-source Python framework for building and orchestrating linear deterministic agentic workflows.

DevenPanchal/simpliflow framework
Open GitHub

Repository containing ready-made example workflows, workflow utilities, and usage material for simpliflow.

DevenPanchal/simpliflow-usage
Open GitHub

Repository for the SimDist framework, including quadruped expert training, simulation data generation, world-model pretraining, deployment, real-world data processing, and finetuning.

CLeARoboticsLab/simdist
Open GitHub

Jericho is presented as an open-source learning environment for human-authored interactive-narrative games.

microsoft/jericho
Open GitHub

TextWorld is presented as a framework for procedural generation and experimentation in text-based games.

microsoft/textworld
Open GitHub

Repository containing supplementary materials for the paper, organized into metrics, heatmaps, and equity-line outputs.

Maciej-13/spxw-options-sizing-paper-appendix
Open GitHub

Official repository for the skfolio Python library, containing the implementation of the portfolio-optimization and risk-management framework described in arXiv:2507.04176.

skfolio/skfolio framework
Open GitHub

Official repository for the paper's SHROOM system experiments, including Data and Experiments directories.

Sharif-SLPL/SE-2024-Task-06-SHROOM
Open GitHub

The repository is linked as the paper's website and contains project materials/figures for Small Language Models: Survey, Measurements, and Insights.

UbiquitousLearning/SLM_Survey
Open GitHub

Official Natural Questions repository linked by the paper for the NQ-Open evaluation benchmark.

google-research-datasets/natural-questions dataset
Open GitHub

Repository containing SMART-LLM task-plan generation and execution scripts, robot definitions and skills, and the benchmark task data used for AI2-THOR evaluation.

SMARTlab-Purdue/SMART-LLM dataset
Open GitHub

Qwen3-Coder repository linked by the paper for the 30B open-weights code-generation model evaluated under NO-SPARK and WITH-SPARK conditions.

QwenLM/Qwen3-Coder
Open GitHub

Repository for the Smoothie label-free LLM routing method.

HazyResearch/smoothie implementation
Open GitHub

Public Python implementation associated with the proposed social recommendation system.

BehafaridMjf/Social-Recommendation-System implementation
Open GitHub

Official Google Research directory containing self-contained prototype notebooks for the Socratic Models applications evaluated in the paper.

google-research/google-research
Open GitHub

Repository containing recovery, relabeling, validation, environment, download, and reproduction scripts for LPQLD on ImageNet-1K and ImageNet-21K-P.

he-y/soft-label-pruning-quantization-for-dataset-distillation
Open GitHub

Modular plug-in framework and CLI for composing target models, attacks, defenses, datasets, and judgers, with released evaluation records and experiment configurations.

datasec-lab/PromptSecurity benchmark
Open GitHub

Repository implementing the Sim2Reason scene-generation, simulation, QA-generation/filtering, RL training, checkpoint, and benchmark-evaluation workflow.

Sim2Reason/Sim2Reason
Open GitHub

llama.cpp version 4501 is the execution backend for all experiments, and its llama-quantize tool produces the Q4_0 and Q4_K_M model variants.

ggerganov/llama.cpp
Open GitHub

Repository for YOLOv8 traffic-density estimation, used as the basis of the Traffic Agent's vehicle detection and congestion-count pipeline.

FarzadNekouee/YOLOv8_Traffic_Density_Estimation
Open GitHub

Repository for the YOLO11 Flare Guard real-time smoke and fire detector used by the Safety Agent.

sayedgamal99/Real-Time-Smoke-Fire-Detection-YOLO11
Open GitHub

Repository containing social-task generation, conversation generation and processing, self-training/RL components, model deployment, automatic evaluation, human evaluation, plotting, and reproduction instructions.

sotopia-lab/sotopia-pi
Open GitHub

Official SPA-Bench repository containing benchmark code, data, framework, model server, pipeline, documentation, and setup files.

ai-agents-2030/SPA-Bench dataset
Open GitHub

Official research repository for Sparse High Rank Adapters, including the paper summary and implementation references.

Qualcomm-AI-research/SHiRA implementation
Open GitHub

Repository associated with the Sparse Logit Sampling paper and Random Sampling KD method; at extraction time the README stated that code would be uploaded after company approvals.

akhilkedia/RandomSamplingKD implementation
Open GitHub

Repository released by the authors for the Sparsing Law study, including code and checkpoints.

thunlp/SparsingLaw
Open GitHub

GitHub Spec Kit documentation for spec-driven development, including the staged specification, planning, and tasking workflow used as the conceptual and procedural base for Spec Kit Agents.

github/spec-kit framework
Open GitHub

Repository containing training and distillation scripts, BigBench Hard and mathematical-benchmark evaluation code, processed-data workflows, prompting notebooks, and notebooks for aligning code-davinci and FlanT5 tokenized outputs.

FranxYao/FlanT5-CoT-Specialization implementation
Open GitHub

GitHub repository associated with the SPEECH method and released resources.

zjunlp/SPEECH dataset
Open GitHub

A curated repository associated with the paper that organizes efficient architecture papers according to the survey's categories.

weigao266/Awesome-Efficient-Arch
Open GitHub

Repository for the Semantic Pyramid Indexing FAISS/Qdrant plug-in introduced and evaluated by the paper.

FastLM/SPI_VecDB framework
Open GitHub

Official code repository for SSAST, including self-supervised pretraining, downstream fine-tuning recipes, SUPERB integration, and released pretrained models.

YuanGongND/ssast
Open GitHub

Mesh TensorFlow MoE implementation cited by the paper as code for its models.

tensorflow/mesh implementation
Open GitHub

Repository for StableToolBench, the paper's stability-oriented benchmark built on ToolBench.

THUNLP-MT/StableToolBench benchmark
Open GitHub

Repository released by the authors containing data-making scripts, training code, safety and reasoning benchmark code, and overrefusal-ablation utilities.

UCSC-VLAA/STAR-1
Open GitHub

Repository explicitly linked by the paper for Star-Agents code. At extraction time, the public repository contains a README stating that the code will be coming soon rather than a released implementation.

CANGLETIAN/Star-Agents
Open GitHub

Official STaRK repository containing benchmark loading utilities, preprocessing code, embeddings helpers, evaluation scripts, configuration, and documentation.

snap-stanford/STaRK dataset
Open GitHub

Official KernelBench repository used by the paper to benchmark runtime of PyTorch baselines and LLM-generated kernels.

ScalingIntelligence/KernelBench dataset
Open GitHub

Repository for the SciBORG manuscript, with setup instructions, framework usage examples, agent construction, and benchmarking guidance tied to the paper release.

chopralab/sciborg_manuscript_repo framework
Open GitHub

Directory containing SciBORG benchmark notebooks and trace notebooks corresponding to the supporting-information experiments.

chopralab/sciborg benchmark
Open GitHub

Repository reported by the paper as the source code for reproducing the pumped-storage DDQN state-representation study.

Fluxons/hydrodam implementation
Open GitHub

Repository containing the processed datasets, subject code, static-analysis implementation, machine-learning pipeline, scripts, outputs, feature selection, tuning, cross-validation, and SHAP analyses used in the paper.

imran9pk/replication-package_method_energy_java
Open GitHub

Repository containing historical changes in S&P 500 constituents, used to restrict backtest positions to stocks that belonged to the index on each date.

fja05680/sp500 dataset
Open GitHub

IMPROVER is an operational probabilistic weather forecast post-processing system used to bias-correct, calibrate, threshold, smooth, and blend forecast products.

metoppv/improver implementation
Open GitHub

Contains data-engineering utilities, synthetic-data generation, the three experimental training flows, SFT, DPO and DTFT configurations, and instructions for reproducing the final StatLLaMA path.

HuangDLab/StatLLaMA implementation
Open GitHub

Official repository for the Steve-Eye paper and project; it contains the project README, paper material, benchmark-result tables, and is linked from the authors' project website.

BAAI-Agents/Steve-Eye benchmark
Open GitHub

Repository for the BESSTIE sentiment and sarcasm classification benchmark for varieties of English.

unswnlp/BESSTIE dataset
Open GitHub

Repository for the InstruSum instruction-controllable summarization dataset referenced and used as an evaluation target in the paper.

yale-nlp/InstruSum dataset
Open GitHub

Repository indicated by the paper as containing all code and supplementary materials used for the SV-LSTM hybrid model study.

aperekhodko/sv_lstm_hybrid_model implementation
Open GitHub

Repository linked by the authors as the location of the curated dataset used for the stock movement and volatility prediction experiments.

hao1zhao/Bigdata23 dataset
Open GitHub

Repository titled MSGCA: Stock Movement Prediction with Multimodal Stable Fusion via Gated Cross-Attention Mechanism, containing code and data folders for the proposed framework.

changzong/MSGCA dataset
Open GitHub

Repository containing code, configuration files, data-preparation scripts, pretraining routines, and trading scripts for STORM.

DVampire/Storm framework
Open GitHub

Official repository for the StoryScope pipeline, configuration, feature taxonomy, development stories, feature assignments, trained XGBoost models, and reproduction scripts.

jenna-russell/storyscope dataset
Open GitHub

Official ECCV 2026 release containing StoryAD-QA annotations, answer keys, evaluation code, and generation and answering prompts; it does not redistribute copyrighted movie media.

SEE-AI-Lab/ECCV2026_StoryTeller_StoryAD_QA
Open GitHub

Code repository for the LLM-agent Cournot competition simulations and experimental workflow.

smojha/collusive-llm-agents implementation
Open GitHub

The referenced Llama 3 model card is associated with Llama-3-8B-Instruct, one of the local open-source models used in the ranking task.

meta-llama/llama3
Open GitHub

Code released for the paper, including modular segmentation-attention and syntactic-attention layers and training scripts for translation, question answering, and natural language inference.

harvardnlp/struct-attn
Open GitHub

Searchat is a local-first semantic search system for AI coding-agent conversations, supporting verbatim, distilled, and cross-layer retrieval over agent transcript histories.

Process-Point-Technologies-Corporation/searchat framework
Open GitHub

Open-source AI hedge fund project whose structured-summary prompt template, JSON action schema, and next-open execution conventions are adapted for the paper's LLM trading-agent backtests.

virattt/ai-hedge-fund framework
Open GitHub

Public reference implementation of Context-Aware Decoding, the decoding method adapted by FinCAD for parametric look-ahead-bias mitigation.

xhan77/context-aware-decoding implementation
Open GitHub

The public repository used to collect and release the Natural Instructions/Super-NaturalInstructions tasks and to provide official cross-task benchmark splits and experiment code.

allenai/natural-instructions dataset
Open GitHub

Open-source SuperHF training code, reward-model training code, PPO-RLHF baselines, experiments, evaluations, and chart-generation resources.

openfeedback/superhf implementation
Open GitHub

Google Research TensorFlow implementation of the SupCon method.

google-research/google-research
Open GitHub

PyTorch implementation of Supervised Contrastive Learning and SimCLR-style training.

HobbitLong/SupContrast
Open GitHub

Repository for the code used in the paper, with directories for the rate, prepare, assess, and pipeline stages.

COMSYS/artifact-evaluation-llm-support system
Open GitHub

Contains benchmark evaluation code, inference utilities, prompts, training code, and links to SURDS/SURDS_eval data used to reproduce the paper's evaluation and post-training workflow.

XiandaGuo/Drive-MLLM
Open GitHub

Companion repository that organizes works on LLM-agent evaluation according to the survey's structure and tracks papers, benchmarks, methodologies, and frameworks.

Asaf-Yehudai/LLM-Agent-Evaluation-Survey
Open GitHub

Framework for evaluating and optimizing agents and models in container environments, discussed as part of emerging standardized cross-environment agent evaluation.

harbor-framework/harbor framework
Open GitHub

LangChain AgentEvals package for evaluating agent trajectories, including trajectory matching and graph-based evaluation.

langchain-ai/agentevals framework
Open GitHub

HAL harness for centralized and reproducible evaluation across agent benchmarks.

princeton-pli/hal-harness
Open GitHub

Repository for SWE-agent, the LM-based agent system that attempts to fix GitHub issues using an agent-computer interface and configurable tools.

SWE-agent/SWE-agent framework
Open GitHub

Repository containing SWE-Bench-CL data, dataset construction scripts, naive and agentic evaluation procedures, and LangGraph/FAISS-based continual-learning agent implementations.

thomasjoshi/agents-never-forget dataset
Open GitHub

Official SWE-bench repository containing benchmark tooling and resources for evaluating language models on real-world GitHub software-engineering issues.

SWE-bench/SWE-bench dataset
Open GitHub

Repository linked by the paper for the SWE-CI benchmark, associated code, and evaluation resources.

SKYLENAGE-AI/SWE-CI dataset
Open GitHub

Repository for the SWE-EVO benchmark, including benchmark materials and evaluation support for coding agents in long-horizon software evolution scenarios.

SWE-EVO/SWE-EVO dataset
Open GitHub

Official repository released for the SWE-Lancer benchmark, public Diamond split, code, and evaluation environment.

openai/SWELancer-Benchmark dataset
Open GitHub

Repository for SWE-QA-Pro, including evaluation materials for direct and agent modes and links to the paper and benchmark release.

TIGER-AI-Lab/SWE-QA-Pro dataset
Open GitHub

Repository containing the open-sourced inference implementation for serving SwiftKV-transformed models.

snowflakedb/arcticinference
Open GitHub

Repository containing the open-sourced training and distillation implementation used to adapt models for SwiftKV.

snowflakedb/arctictraining
Open GitHub

The paper states that JAX code for Switch Transformer and all model checkpoints are available in T5X.

google-research/t5x implementation
Open GitHub

The paper points to the Mesh TensorFlow MoE implementation as TensorFlow code for Switch Transformer.

tensorflow/mesh implementation
Open GitHub

Official inference repository for FLUX.1 models, including FLUX.1-dev and inpainting/editing variants.

black-forest-labs/flux
Open GitHub

Repository containing experiment code for SynthSAEBench SAE architecture evaluations.

decoderesearch/synth-sae-bench-experiments benchmark
Open GitHub

Official repository for the T-Eval evaluation harness, benchmark protocols, test scripts, data access, and model-evaluation workflow.

open-compass/T-Eval dataset
Open GitHub

Codebase provided by the authors for reproducing TABCF experiments.

Panagiotou/TABCF implementation
Open GitHub

The paper states that this repository contains the reproducibility code for the benchmark of tabular classification methods.

machinelearningnuremberg/TabularStudy benchmark
Open GitHub

Repository containing the TAGAL implementations, prompt templates, dataset loaders, competitor implementations, generated examples, evaluation outputs, logs, and table-generation sources.

bronval/TAGAL-Tabular-Data-Generation-with-LLMs implementation
Open GitHub

Official TaskBench code and dataset directory within the Microsoft JARVIS repository.

microsoft/JARVIS dataset
Open GitHub

A JSON file encoding TDD principles as governance objects with bibliographic grounding, human-oriented intent, AI-native interpretation, operational constraints, and anti-patterns.

shahbazsiddeeq/TDD-manifesto implementation
Open GitHub

Repository linked by the paper for data and code supporting Self-Critique Fine-Tuning and Reinforcement Learning with Effective Reflection Rewards.

wanghanbinpanda/SCFT
Open GitHub

Open-source Python framework extending SMAC with meta-learning and ensemble learning for pipeline automation.

automl/auto-sklearn framework
Open GitHub

Open-source Python implementation of sequential model-based algorithm configuration for hyperparameter optimization.

automl/SMAC3 framework
Open GitHub

Open-source Python Auto-Pipeline tool using genetic programming to optimize tree-structured machine learning pipelines.

EpistasisLab/tpot framework
Open GitHub

Open-source Python framework for automated feature generation from relational datasets.

Featuretools/featuretools framework
Open GitHub

Open-source Python tool for hyperparameter optimization.

hyperopt/hyperopt framework
Open GitHub

Open-source Keras-based framework for searching deep network architectures using Bayesian optimization and network morphism.

keras-team/autokeras framework
Open GitHub

Microsoft open-source toolkit for neural architecture search and hyperparameter tuning across local or cloud execution environments.

microsoft/nni framework
Open GitHub

Open-source TensorFlow framework for automatically learning neural network architectures and ensembles.

tensorflow/adanet framework
Open GitHub

Repository released by the authors for the IN3 benchmark, Mistral-Interact training/inference code, interaction data, and evaluation scripts.

HBX-hbx/Mistral-Interact dataset
Open GitHub

Open LLaMA is the publicly available 13B model used in the paper's instruction-based fine-tuning experiments.

openlm-research/open_llama implementation
Open GitHub

Repository containing the TGN model, modules, data preprocessing utilities, self-supervised link-prediction training, supervised dynamic node-classification training, baseline configurations, and ablation commands.

twitter-research/tgn implementation
Open GitHub

Official TemporalBench repository containing evaluation scripts, scoring code, model-evaluation examples, and links to the released dataset and leaderboard.

mu-cai/TemporalBench benchmark
Open GitHub

Repository for the TensorFlow.js framework introduced and described by the paper.

tensorflow/tfjs
Open GitHub

Repository for Tevatron; the current project has evolved beyond the original release and documents access to the v1 branch for original Tevatron features.

texttron/tevatron
Open GitHub

Repository containing TextAtari agent/environment code, translators from Gym-style environments to natural language, prompt/decider modules, manuals, language trajectories, and visualization assets.

Lww007/Text-Atari-Agents system
Open GitHub

Python implementation of the three-stage TextReg pipeline, including RuleBank, gradient purification, semantic edit regularization, the guided optimizer, example execution scripts, and paper-aligned metrics.

luchengfu6/TextReg framework
Open GitHub

Repository accompanying the paper, with runnable TEP pipelines, configurations, benchmark support, and implementation code for deep compound AI system optimization.

MinghuiChen43/TEP framework
Open GitHub

Repository for the QA-FEEDBACK, Longformer reward-model, and T5/PPO experimental workflow used to study the reward-model accuracy paradox.

EIT-NLP/AccuracyParadox-RLHF implementation
Open GitHub

Repository associated with the AI Cosmologist paper, described by the paper as containing code and experimental data and by the repository README as containing configuration files, examples, and best AI-generated code for the Galaxy Zoo and Quijote demonstrations.

adammoss/aicosmologist system
Open GitHub

Contains the three generated submissions and associated experiment, review, and audit materials for compositional regularization, label noise and calibration, and pest detection.

SakanaAI/AI-Scientist-ICLR2025-Workshop-Experiment
Open GitHub

Contains the template-free autonomous research pipeline, ideation and experimentation code, best-first tree-search configuration, prompts, launch scripts, and implementation documentation.

SakanaAI/AI-Scientist-v2
Open GitHub

Repository providing the ICLR 2022 OpenReview data used to evaluate the automated reviewer.

fedebotu/ICLR22-OpenReviewData
Open GitHub

Official repository for The AI Scientist framework, prompts, code templates, generated papers, run files, logs, and automated-review implementation.

SakanaAI/AI-Scientist
Open GitHub

Repository linked by the paper for LLM uncertainty decomposition, with folders and scripts related to input uncertainty, decoding uncertainty, model uncertainty, data, models, utilities, and uncertainty scoring.

aditya-taparia/LLM-Uncertainty implementation
Open GitHub

Repository released by the authors for the agent-to-agent negotiation and transaction benchmark code and data.

ShenzheZhu/A2A-NT benchmark
Open GitHub

Official BRAVO challenge repository containing the evaluation toolkit, baselines, submission encoding utilities, benchmark instructions, dataset access information, and challenge results.

valeoai/bravo_challenge benchmark
Open GitHub

ParlAI project directory containing the CRINGE loss implementation and iterative-training support code, including cringe_loss.py, teachers.py, and safety_filter_world_logs.py.

facebookresearch/ParlAI
Open GitHub

Official repository for the paper, containing generation code, data-formatting utilities, metric-related files, and the Mechanical Turk template used in evaluation.

ari-holtzman/degen
Open GitHub

Meta Research repository associated with the paper's MAE?WSP pre-pretraining approach.

facebookresearch/maws implementation
Open GitHub

Code repository for the paper's controlled assessment pipeline, mitigation strategies, and fidelity-resistance evaluation.

ASTRAL-Group/BDC-mitigation-assessment
Open GitHub

Repository containing scripts and pre-computed experimental results for the paper's probing, automatic expert labeling, trigger-target attribution, clustering, specialization analysis, and figure reproduction.

jerryy33/MoE_analysis
Open GitHub

Code repository linked by the paper for the controlled case studies evaluating attribution methods under the two constructed scenarios.

AmTuTi1999/FMIBAMGTSE
Open GitHub

Datatrove is the large-scale data-processing library used to build FineWeb; the paper links a FineWeb reproduction script in this repository.

huggingface/datatrove
Open GitHub

Official repository for Calibration-Aware On-Policy Distillation, the method introduced and evaluated by the paper.

SalesforceAIResearch/CaOPD implementation
Open GitHub

Public implementation of Progressive Transformers retrained and evaluated to provide a BLEU reference scale.

BenSaunders27/ProgressiveTransformersSLP
Open GitHub

Implementation repository for the sign-pose VAE variants introduced and evaluated in the paper.

GFaure9/SignPoseVAE implementation
Open GitHub

Public implementation of Sign-IDD retrained and evaluated as a non-latent diffusion reference baseline.

NaVi-start/Sign-IDD
Open GitHub

Repository containing the paper's transparency materials, including SP-1 AI-usage summary, SP-2 navigation index, SP-3 documentation-adequacy account, SP-4 process documentation, and SP-5 development records.

MicheleLoi/JPEP implementation
Open GitHub

A framework for training or evaluating agentic systems with reinforcement learning.

Agent-One-Lab/AgentFly framework
Open GitHub

A Microsoft framework listed as part of the agentic RL framework ecosystem.

microsoft/agent-lightning framework
Open GitHub

An open-source framework for scaling LLM reinforcement learning.

NovaSky-AI/SkyRL framework
Open GitHub

A repository for verifiable environments used in LLM reinforcement learning.

PrimeIntellect-ai/verifiers framework
Open GitHub

A benchmark for evaluating agents on real software engineering issues.

swe-bench/SWE-bench benchmark
Open GitHub

A framework for tool-use reinforcement learning with LLM agents.

TIGER-AI-Lab/verl-tool framework
Open GitHub

A web environment benchmark used to evaluate autonomous agents on browser tasks.

web-arena-x/webarena benchmark
Open GitHub

Repository for the paper's released data, including generated Code Llama outputs used to support follow-on work on ranking LLM code-generation candidates.

slp-rl/budget-realloc dataset
Open GitHub

Repository accompanying the paper, containing SPR task materials, generated research outputs, case studies for Agent Laboratory and The AI Scientist v2, and pitfall-detection code.

niharshah/AIScientistPitfalls dataset
Open GitHub

Open-source implementation of The AI Scientist-v2, one of the two AI scientist systems evaluated in the paper.

SakanaAI/AI-Scientist-v2 system
Open GitHub

Open-source implementation of Agent Laboratory, one of the two AI scientist systems evaluated in the paper.

SamuelSchmidgall/AgentLaboratory system
Open GitHub

Official Google DeepMind repository containing NarrativeQA document metadata, Wikipedia summaries, question-answer pairs, and story-download/verification scripts.

google-deepmind/narrativeqa dataset
Open GitHub

Repository released by the authors for obtaining and preprocessing decaNLP datasets, training and evaluating models, reproducing experiments, and tracking decaScore progress.

salesforce/decaNLP
Open GitHub

AI-Descartes combines symbolic regression with formal reasoning to evaluate whether candidate scientific formulas are derivable from background axioms.

IBM/AI-Descartes
Open GitHub

AI-Hilbert integrates experimental data and polynomial background axioms during candidate-law generation and can output exact or approximate derivability certificates.

IBM/AI-Hilbert
Open GitHub

Official repository releasing generated stories and the crowdworker and in-house human-evaluation results used by the paper.

ZhuohanX/TheNextChapter
Open GitHub

NeuralMagic MLPerf Inference v2.1 submission referenced by the paper for the compressed BERTLARGE and MobileBERT deployment results.

neuralmagic/mlperf_inference_results_v2.1
Open GitHub

SparseML research directory containing oBERT scripts, recipes, tutorials, checkpoints links, and the OBS pruning implementation used to reproduce paper results.

neuralmagic/sparseml
Open GitHub

Meta Llama Guard 2 model-card repository path for the safeguard model used to classify whether generated outputs violate predefined safety categories.

meta-llama/PurpleLlama implementation
Open GitHub

Official Google Research codebase for reproducing the paper's prompt-tuning experiments, with training configurations, released prompts, and T5.1.1 LM-adapted checkpoints.

google-research/prompt-tuning
Open GitHub

T5 code, evaluation metrics, preprocessing routines, and checkpoint index used by the paper for its base models, data preparation, and reproducibility.

google-research/text-to-text-transfer-transformer
Open GitHub

Resources and out-of-domain development data for the MRQA 2019 shared task used in the paper's zero-shot question-answering transfer experiments.

mrqa/MRQA-Shared-Task-2019
Open GitHub

Repository containing the training, reward, evaluation, dataset-conversion, and inference code used for the paper's legal machine translation experiments.

aixiuxiuxiu/Legal-MT-SFT-RL
Open GitHub

Repository associated with the Test-Driven AI Agent Definition paper, containing code or benchmark artifacts for compiling tool-using agents from behavioral specifications.

f-labs-io/tdad-paper-code benchmark
Open GitHub

GitHub Spec Kit is an open-source toolkit for specification-driven development using commands such as /speckit.constitution, /speckit.specify, /speckit.plan, /speckit.tasks, /speckit.analyze, and /speckit.implement.

github/spec-kit framework
Open GitHub

Official repository containing code and supporting materials for automated literature collection, deduplication, filtering, review assistance, and the paper's experiments.

trigaten/The_Prompt_Report
Open GitHub

Companion code and artifact repository containing analysis scripts, prompt templates, query lists, judge labels, numerical result JSONs, probe configurations, leakage diagnostics, Procrustes tests, and steering analyses.

amanmehta-maniac/refusal-residue-release
Open GitHub

Anthropic's official Skill-Creator used as one of the three meta-skill creators that generate treatment libraries from shared failure signals.

anthropics/skills
Open GitHub

OpenAI/Codex official Skill-Creator used as one of the three meta-skill creators that generate treatment libraries from shared failure signals.

openai/skills
Open GitHub

Sentient skill-creator repository corresponding to the harness-agnostic meta-skill creator written for the study.

sentient-agi/meta-skill-creator
Open GitHub

A GitHub repository collecting papers related to LLM-based agents, linked by the survey as a related-papers resource.

WooooDyy/LLM-Agent-Paper-List
Open GitHub

Public repository for the AIDev dataset, schema/CSV files, replication materials, and example notebooks supporting analyses of autonomous coding agents in GitHub PR workflows.

SAILResearch/AI_Teammates_in_SE3 dataset
Open GitHub

Repository containing SHS source code, documentation, examples, test cases, anonymized participant responses, data dictionary, R and Python analysis scripts, questionnaires, and study protocol materials.

human-centered-ai-lab/system-hallucination-scale
Open GitHub

Repository linked by the paper for the reproduction work associated with the study.

AIReproducibility2018 implementation
Open GitHub

Repository containing released evaluation experiments and artifacts for baseline agents and model configurations on TheAgentCompany.

TheAgentCompany/experiments
Open GitHub

Repository for the TheAgentCompany benchmark environment, tasks, data, and evaluation infrastructure introduced by the paper.

TheAgentCompany/TheAgentCompany dataset
Open GitHub

Repository implementing the Theorem-of-Thought multi-agent reasoning pipeline, including baselines, data, and the core ToTh code.

KurbanIntelligenceLab/theorem-of-thought
Open GitHub

Repository containing the paper's multi-LLM collaboration and ToM reference implementation, including the gym-dragon environment and scripts for LLM-agent experiments.

romanlee6/multi_LLM_comm
Open GitHub

Official code repository for DT-Mem, including training and fine-tuning code for the memory-augmented Decision Transformer.

luciferkonn/DT_Mem
Open GitHub

A project associated with the survey's core-competency test framework for LLM evaluation.

HITSCIR-DT-Code/Core-Competency-Test-for-the-Evaluation-of-LLMs dataset
Open GitHub

Repository for the TileMix implementation introduced and evaluated by the paper.

HanzhiZhang-Ulrica/TileMix
Open GitHub

Repository containing the DeepFund code used to implement the live fund-investment benchmark, agent workflow, prompts, and evaluation system.

HKUSTDial/DeepFund system
Open GitHub

The official Time-MoE repository associated with the paper's architecture and released resources.

Time-MoE/Time-MoE dataset
Open GitHub

Repository containing the TimeArena environment, resources, evaluation scripts, metric code, prompts, and oracle-calculation code used to run the benchmark.

ykzhang721/TimeArena benchmark
Open GitHub

Repository for the TimeBlind benchmark with setup instructions, data format, evaluation code, and metric implementation.

Baiqi-Li/TimeBlind dataset
Open GitHub

Repository containing code, data-processing materials, experiment scripts, results, and plotting notebooks for reproducing the paper's analyses.

felipemaiapolo/efficbench implementation
Open GitHub

Repository containing the tinyBenchmarks Python package, demos, tutorials, and links to tiny datasets for estimating LLM performance from curated small benchmark subsets.

felipemaiapolo/tinyBenchmarks package
Open GitHub

Official repository for the TLOB paper, including code folders for models, preprocessing, data handling, configuration, training scripts, requirements, and backtesting script.

LeonardoBerti00/TLOB benchmark
Open GitHub

Official repository containing TofuEval development/test identifiers and released annotations for factual consistency, written explanations, hallucination types, completeness key points, and topic categorization.

amazon-science/tofueval dataset
Open GitHub

Repository implementing TokUR's low-rank Bayesian weight perturbation, uncertainty evaluation, dataset preparation, and test-time-scaling experiments.

Wang-ML-Lab/TokUR
Open GitHub

BMTools integrates the paper's evaluated tools and provides an open-source platform for extending foundation models with APIs and for building and sharing tool plugins.

OpenBMB/BMTools dataset
Open GitHub

Repository for ToolRoCo, a multi-turn tool-using LLM benchmark for collaborative robotic tasks with Cabinet, PackGrocery, and Sort tasks and four cooperation paradigms.

ColaZhang22/Tool-Roco benchmark
Open GitHub

Official repository containing the ToolAlpaca data, prompts, multi-agent generation code, training scripts, evaluation code, and recorded evaluation outputs.

tangqiaoyu/ToolAlpaca dataset
Open GitHub

Hosts the ToolCAD project website and paper-facing artifact page.

gongyifeiisme/toolcad-project system
Open GitHub

Repository currently titled Ziqiao-git/C-World but with README content for ToolGym. It describes ToolGym as an open-world tool-using environment built on 5,571 tools across 204 applications and includes code folders for task creation, tool retrieval, state controller, runtime, and evaluation.

Ziqiao-git/C-World dataset
Open GitHub

Repository for ToolBench/ToolLLM artifacts, including code, trained models, and demo released by the authors.

OpenBMB/ToolBench dataset
Open GitHub

Repository for the ToolMisuseBench benchmark implementation, generator, evaluator, and experiment reproduction workflow.

akgitrepos/toolmisusebench benchmark
Open GitHub

Repository for ToolPRMBench, the benchmark and associated code/data for evaluating process reward models in tool-using agents.

David-Li0406/ToolPRMBench dataset
Open GitHub

Detectron2 is listed among object detection and image segmentation models/frameworks in the TorchTraceAP application table.

facebookresearch/detectron2 framework
Open GitHub

PyTorch Holistic Trace Analysis package referenced as the source of profiling metrics and rule-based trace-event runtime outlier analysis.

facebookresearch/HolisticTraceAnalysis package
Open GitHub

Repository containing the Touchdown corpus splits, graph files, navigation/SDR data structures, helper code, and license information.

lil-lab/touchdown dataset
Open GitHub

Open-source implementation of Self-Supervised Contrastive Pre-Training for Time Series via Time-Frequency Consistency.

mims-harvard/TFC-pretraining implementation
Open GitHub

GitHub repository for the Light Aircraft Game benchmark, which the paper includes as an application simulator in its benchmark characterization table.

liuqh16/CloseAirCombat benchmark
Open GitHub

The BosqueLanguage GitHub organisation is identified by the paper as the public location for experimental versions of the AISE-related systems, including the Bosque ecosystem components discussed in the paper.

BosqueLanguage framework
Open GitHub

Repository containing the implementation, experiment configurations, synthetic-ring generation procedure, and supporting evaluation artifacts for the auditable fraud-detection and agentic-investigation pipeline studied in the paper.

rahil1303/auditable-fraud-investigation
Open GitHub

Repository containing X-FM code and pre-trained models.

zhangxinsong-nlp/XFM implementation
Open GitHub

Repository containing code for survival-model calibration methods including CiPOT and CSD-style calibration, plus experiment reproduction resources.

shi-ang/MakeSurvivalCalibratedAgain implementation
Open GitHub

A GitHub repository linked by the paper/project page that curates efficient-agent papers and resources corresponding to the survey taxonomy.

yxf203/Awesome-Efficient-Agents
Open GitHub

Repository linked by the paper for code supporting its synthetic-GMM simulation and KL-divergence analysis.

ZyGan1999/Towards-a-Theoretical-Understanding-of-Synthetic-Data-in-LLM-Post-Training
Open GitHub

Repository containing prompts and queries used in the experiments for the LLM-powered security alert investigation workflow.

Rub3cula/CyberHunt2025 system
Open GitHub

A maintained collection of papers, methods, benchmarks, and resources associated with RAG-reasoning and agentic deep-research systems.

DavidZWZ/Awesome-RAG-Reasoning
Open GitHub

Apollo is an open autonomous driving platform used as the representative automated-driving system context for the safety-requirements derivation task.

ApolloAuto/apollo system
Open GitHub

Official repository for the TAT-DQA dataset and benchmark introduced by the paper.

NExTplusplus/TAT-DQA dataset
Open GitHub

Repository containing processed evaluation datasets and code for WPQ, Local Order Quiz, Token Overlap, Canonical Order, and Min-K% experiments.

vsamuel2003/data-contamination
Open GitHub

Repository for the ACL 2026 survey that records system-aware, serving-time, KV-centric optimization papers and organizes them using the paper's temporal, spatial, and structural taxonomy.

jjiantong/Awesome-KV-Cache-Optimization
Open GitHub

Code repository implementing the paper's unified MoE compression framework and proposed Expert Trimming methods.

CASE-Lab-UMD/Unified-MoE-Compression framework
Open GitHub

ROS framework for embodied intelligence applications using robot API configuration and LLM calls.

Auromix/ROS-LLM framework
Open GitHub

ROS 2 command-line interface extension with LLM support.

fujitatomoya/ros2ai system
Open GitHub

Stretch AI system orchestrating skills for language-directed mobile manipulation.

hello-robot/stretch_ai implementation
Open GitHub

MCP server connecting AI assistants to installed ROS 2 applications and system operations.

lpigeon/ros-mcp-server implementation
Open GitHub

Project integrating llama.cpp with ROS 2 to enable LLM inference.

mgonzs13/llama_ros implementation
Open GitHub

Tool using LLMs to generate ROS codebases from high-level descriptions.

RoboCoachTechnologies/ROScribe system
Open GitHub

ROS/ROS2 MCP server enabling natural-language commands and monitoring of robot states and sensor data.

robotmcp/ros-mcp-server
Open GitHub

ROS-MCP project included among the paper's representative ROS/MCP integrations.

Yutarop/ros-mcp implementation
Open GitHub

Repository containing the implemented prototype, generated code, and execution traces for the LLM workflow generation experiments.

dos-group/LLMWorkflowGenerator system
Open GitHub

Repository linked by the authors as the paper list for the survey on reasoning in large language models.

jeffhj/LM-reasoning
Open GitHub

Haystack is the framework used to implement the paper's RAG workload with retrieval and question-answering pipeline behavior.

deepset-ai/haystack framework
Open GitHub

GitHub repository stated by the paper as the available source code for the multimodal financial forecasting work.

sarthak-12/thesis-dsaa framework
Open GitHub

Repository containing TraceCoder source code, position-key implementation, relational storage and viewer components, experiment problems, run scripts, and result artifacts.

devfitcs/TraceCoder
Open GitHub

DeepMarket is the official open-source Python framework for LOB market simulation with deep learning. It contains TRADES and CGAN implementations/checkpoints, ABIDES-based simulation components, evaluation utilities, and the TRADES-LOB synthetic dataset.

LeonardoBerti00/DeepMarket dataset
Open GitHub

Repository for TradeTrap, the paper's system-level stress-testing framework for LLM-based autonomous trading agents.

Yanlewen/TradeTrap framework
Open GitHub

Repository for TradingAgents, the multi-agent LLM financial trading framework introduced and evaluated in the paper.

TauricResearch/TradingAgents framework
Open GitHub

Code repository for the LoRAM and QLoRAM training, recovery, alignment, and evaluation workflow introduced by the paper.

junzhang-zj/LoRAM
Open GitHub

Repository associated with the paper's helpfulness preference data for training and evaluating helpful and harmless assistants.

anthropics/hh-rlhf dataset
Open GitHub

Repository for BIG-bench, one of the principal benchmark suites used to compare Chinchilla with Gopher across diverse language-model capabilities.

google/BIG-bench benchmark
Open GitHub

Original research code for DDPO and the RWR baselines used for the paper's diffusion reinforcement-learning experiments.

jannerm/ddpo
Open GitHub

Official PyTorch implementation of DDPO with GPU and LoRA support, linked as an update from the paper's project page.

kvablack/ddpo-pytorch implementation
Open GitHub

Repository containing released model samples for sampling-based NLP evaluations reported in the paper.

openai/following-instructions-human-feedback
Open GitHub

Repository containing reference code for ILF experiments, refinement scoring, reward models, evaluation scripts, and links to the released SLF5K and finetuning datasets; the authors note that excluded data-generation and cleaning steps mean it is not a ready-to-run reproduction package.

JeremyAlain/imitation_learning_from_language_feedback implementation
Open GitHub

Contains Sandbox social-simulation code, released interaction data, Stable Alignment training code, and links to released socially aligned model checkpoints.

agi-templar/Stable-Alignment
Open GitHub

Repository accompanying the paper's representation-based cut-statistic training-subset selector and experimental pipeline.

hunterlang/weaksup-subset-selection
Open GitHub

Repository containing GSM8K train/test data, calculation annotations, example model solutions, and illustrative calculator/model code.

openai/grade-school-math dataset
Open GitHub

Repository for the TEAL method and sparse inference implementation introduced in the paper.

FasterDecoding/TEAL implementation
Open GitHub

Repository for TRAJECT-Bench, including public data, tool definitions, query generation materials, and evaluation scripts for model and ReAct-style agentic tool-use evaluation.

PengfeiHePower/TRAJECT-Bench dataset
Open GitHub

GitHub data source cited for the Stochastic Block Model dynamic graph benchmark used in the experiments.

IBM/EvolveGCN dataset
Open GitHub

Official ROLAND repository used by the authors to run the ROLAND baseline five times for MAP/MRR comparison.

snap-stanford/roland implementation
Open GitHub

Author-provided code repository for the experiments validating the theoretical critical-point structures of trained linear transformers.

chengxiang/LinearTransformer
Open GitHub

Repository containing Python code, prompt files, forex price data, annotated sentiment data, prediction files, and scripts for reproducing prompt runs and comparative results.

giorgosfatouros/Financial-Sentiment-Analysis-with-ChatGPT dataset
Open GitHub

Repository released by the authors containing code, generated translations, and human quality assessments for the quality-aware cascaded translation system.

deep-spin/translate-smart implementation
Open GitHub

Official repository for TravelPlanner containing agent runners, database assets, tools, postprocessing, evaluation scripts, and benchmark usage instructions.

OSU-NLP-Group/TravelPlanner benchmark
Open GitHub

Official implementation and experiment code for the Treble Counterfactual VLM intervention and its POPE and MMHal-Bench evaluations.

TREE985/Treble-Counterfactual-VLMs
Open GitHub

Repository containing Tree of Thoughts code, task implementations, prompts, and logged experimental trajectories.

princeton-nlp/tree-of-thought-llm framework
Open GitHub

Official CAGE Challenge 4 repository for the multi-agent cyber-defense environment used as one of Trident's two main training and evaluation testbeds.

cage-challenge/cage-challenge-4
Open GitHub

Official CyberWheel repository for the autonomous cyber-defense simulation environment used to train and evaluate Trident's adaptive red agent.

ORNL/cyberwheel
Open GitHub

Official AG2 example repository from which the paper draws four representative MAS applications for case-study evaluation.

ag2ai/build-with-ag2
Open GitHub

Open-source implementation repository for the TrinityGuard framework introduced by the paper.

AI45Lab/TrinityGuard framework
Open GitHub

Repository for TrustAgent code, safety regulations, assets/data, and experiment-running instructions.

agiresearch/TrustAgent dataset
Open GitHub

Repository containing SFT, RLAIF training, preference-data generation, evaluation code, dataset preparation guidance, and links to model checkpoints/datasets.

yonseivnl/vlm-rlaif implementation
Open GitHub

Official repository for TVLT with model code, pretraining and finetuning scripts, demos, setup documentation, and links to released checkpoints.

zinengtang/TVLT
Open GitHub

A custom library referenced by the paper as the interface through which generated Python code controlled the Boston Dynamics Spot robot.

sheepskins/spottyai system
Open GitHub

GitHub repository for the Electricity Transformer Dataset used as ETTh1 and ETTm1 benchmark data in the experiments.

zhouhaoyi/ETDataset dataset
Open GitHub

Repository for a research demonstration of the Lean-Agent Protocol, including a frontend, FastAPI orchestrator, Lean worker, policy environment, audit log, and natural-language-to-Lean/back-translation workflow.

arkanemystic/lean-agent-protocol framework
Open GitHub

Python implementation of UAC, dataset preprocessing, ID and OOD training scripts, and temperature-scaling, entropy-maximization, and Laplace baselines.

Schindler-EPFL-Lab/UAC
Open GitHub

Repository containing the authors' implementation of UMoE, the shared-expert architecture that unifies attention-MoE and FFN-MoE modules.

ysngki/UMoE system
Open GitHub

Author-provided code repository for the unbiased-alignment methods introduced in the paper.

cswjl/unbiased-alignment
Open GitHub

Official Yandex Research repository containing implementations of experiments from this ICLR 2021 paper and the related proxy Dirichlet distillation work.

yandex-research/proxy-dirichlet-distillation
Open GitHub

Meta Llama 3 model documentation for the open-weight model family; the paper evaluates LLaMA-3-8B and cites this GitHub model card.

meta-llama/llama3
Open GitHub

TruthTorchLM is the library the paper states it uses to compare representative uncertainty and truthfulness methods in Section 6.

Ybakman/TruthTorchLM
Open GitHub

Codebase extending tau^2-bench with runtime token-level UQ tracking, trajectory aggregation, uncertainty evaluation, and observation-UQ scoring.

deeplearning-wisc/agentuq
Open GitHub

Repository associated with the uncertainty-manipulation attack and confidentiality-preserving audit protocol.

cleverhans-lab/confidential-guardian system
Open GitHub

Repository associated with training-dynamics-based selective classification experiments.

cleverhans-lab/sc implementation
Open GitHub

Repository associated with experiments and analysis for decomposing the selective-classification gap.

cleverhans-lab/sc-gap implementation
Open GitHub

Repository used to reproduce or support the private selective-classification experiments.

cleverhans-lab/selective-classification implementation
Open GitHub

The vLLM repository provides the production inference engine and fused MoE execution pipeline into which the authors integrate their activation-sparse routed-expert code path.

vllm-project/vllm framework
Open GitHub

Paper-associated GitHub path identified by the authors as the location of the study's code and data.

CyberScienceLab/Our-Papers
Open GitHub

Official repository containing scripts for the paper's first- and second-order experiments across task labels, preferences, instructions, simulations, and free-form text.

minnesotanlp/artifacts-of-llmgendata
Open GitHub

Repository accompanying the paper's controlled bias-generation, downstream evaluation, analysis, and mitigation experiments.

MiaomiaoLi2/bias-inheritance
Open GitHub

Open-source package for replicating experiments, with raw trajectories hosted on Hugging Face, analysis scripts, and data referenced in the paper.

ARiSE-Lab/understanding-apr-agents
Open GitHub

Official SWE-bench experiments repository used to retrieve public agent logs, trajectories, and patch diffs for studied configurations.

swe-bench/experiments
Open GitHub

Repository containing metadata, analysis data, tools, and code for the cryptocoin correlation analysis.

quapsale/cryptoanalytics dataset
Open GitHub

Official Pythia repository for the model suite whose released checkpoints and reported downstream performance are analyzed as an external validation of the paper's loss-performance relationship.

EleutherAI/pythia
Open GitHub

Official repository containing ITERATER datasets, preprocessing code, intent-classification code, revision-model training and evaluation code, and demonstration materials.

vipulraheja/IteraTeR dataset
Open GitHub

Repository for MAFBench, the unified benchmark suite introduced by the paper for controlled evaluation of multi-agent LLM frameworks.

CoDS-GCS/MAFBench benchmark
Open GitHub

Repository for ORCA, described as a step toward automating multi-agent system construction from high-level task descriptions using empirical benchmark evidence and cost-aware execution models.

CoDS-GCS/ORCA system
Open GitHub

Public Python repository containing data, source code, scripts, tests, results, analysis outputs, and paper figures for the Logic-in-LLMs study.

XAheli/Logic-in-LLMs dataset
Open GitHub

Repository released with the paper containing linguistic-rewrite data and resources for response generation, evaluation, and intervention experiments.

nobody294/linguistic_rule_triggers dataset
Open GitHub

Repository for the LLM-ification of CHI review, including sampled CHI papers, qualitative codes, metadata, taxonomy images, and supplementary materials.

rrrrrrockpang/llm-chi dataset
Open GitHub

Open-source implementation of LIME used by the authors to produce local feature-based explanations for the income prediction and biography classification tasks.

marcotcr/lime package
Open GitHub

Repository for the AndroidArena environment, benchmark task files, agent runners, replay/evaluation code, setup scripts, and related configurations.

AndroidArenaAgent/AndroidArena benchmark
Open GitHub

Microsoft's UniLM repository contains the implementation and released pre-trained models for the unified language model introduced in the paper.

microsoft/unilm implementation
Open GitHub

Official DeepMind repository containing the experimental curves and hyperparameters used to derive the paper's main scaling-law results, plus an example notebook for loading the data.

google-deepmind/scaling_laws_for_routing
Open GitHub

Source code for USEagent, the unified software-engineering agent architecture evaluated in the paper.

nus-apr/USEagent framework
Open GitHub

Source code for USEbench, the unified software-engineering benchmark used to evaluate USEagent and baselines.

nus-apr/USEbench dataset
Open GitHub

Official implementation of CROSS for the paper, including model code, utilities for LLM temporal-chain embeddings, training scripts for temporal link prediction, logs, and dataset-preparation instructions.

SiweiPro/CROSS framework
Open GitHub

Official implementation containing SFT, DP fine-tuning, DPO, UnDial-based unlearning and regularization, true-prefix attacks, and task-evaluation scripts.

martonszep/llm-pii-leak
Open GitHub

Repository released by the authors for the paper's universal and transferable LLM attack method and associated AdvBench artifacts.

llm-attacks/llm-attacks
Open GitHub

Repository collecting materials related to data assessment and selection for language-model instruction tuning.

yuleiqin/fantastic-data-engineering
Open GitHub

Provides data-preparation, SynDiff training and translation, tutorship, adaptation, nnU-Net inference, and evaluation workflows for out-of-distribution microscopy segmentation.

axondeepseg/AxonDeepSynth
Open GitHub

Repository for the source code used in the paper's statistical-significance benchmark of online regression over multiple datasets.

mabushaera/Online-Regression-Statistical-Significance benchmark
Open GitHub

Repository for the TextFusionHTS framework and experiments reported by the paper.

xinzzzhou/TextFusionHTS framework
Open GitHub

Official project repository linked by the paper for the UProp method.

jinhaoduan/UProp implementation
Open GitHub

Official AgentBench codebase used for the Operating System agent evaluation.

THUDM/AgentBench benchmark
Open GitHub

Contains UrbanKGent inference code for RTE/KGC, baselines, prompts, sample SFT data, utilities, model-serving scripts, and LoRA fine-tuning code.

usail-hkust/UrbanKGent
Open GitHub

Official Traversaal AI repository for UrduBench containing translated benchmark resources and evaluation notebooks.

traversaal-ai/urdubench_leaderboard dataset
Open GitHub

Official implementation of the precedent-based reaction-plausibility evaluator used as URSA's automated Solv-2 component.

insilicomedicine/ChemCensor
Open GitHub

Official implementation of URSA, including route validation, building-block checks, collapsed route variants, ChemCensor scoring, and dataset-level Solv-N metrics.

insilicomedicine/URSA
Open GitHub

Repository accompanying USF-MAE with pretraining code, preprocessing notebooks, pretrained checkpoints, figures, and access links for OpenUS-46.

Yusufii9/USF-MAE dataset
Open GitHub

Code repository for D3PO, the paper's reward-model-free direct preference fine-tuning method for diffusion models.

yk7333/D3PO implementation
Open GitHub

Repository containing code for Using Non-Expert Data to Robustify Imitation Learning via Offline Reinforcement Learning, including simulation pipeline, scripts, installation instructions, and training/evaluation commands.

UWRobotLearning/RISE implementation
Open GitHub

Repository containing code for cryptocurrency price prediction using RNN-based models and comparison of LSTM, GRU, and Bi-LSTM methods.

shamima08/Cryptocurrency-Price-Prediction-using-RNN implementation
Open GitHub

Repository for V-Rubrics with data conversion, SFT and GRPO training scripts, evaluation code, environments, tests, and release tooling.

shulin16/v-rubrics
Open GitHub

Stable Baselines3 is cited as the implementation source for state-of-the-art baseline algorithms such as SAC, TD3, and PPO used in the experiments.

DLR-RM/stable-baselines3 package
Open GitHub

Codebase for the variational-quantum-circuit DDPG/DQN portfolio agents, baselines, training, evaluation, and reproducibility workflow introduced in the paper.

VincentGurgul/qrl-dpo-public framework
Open GitHub

Microsoft's research repository contains a dedicated VATLM directory with training and model code and links to released checkpoints.

microsoft/SpeechT5
Open GitHub

Flask repository used to illustrate semantic, Louvain, and label-propagation clustering over source files.

pallets/flask
Open GitHub

Python Poetry repository used to demonstrate graph construction, object statistics, issue-driven retrieval, and top-k results.

python-poetry/poetry
Open GitHub

Public repository implementing the VeinCast architecture, training curriculum, evaluation, and inference workflow introduced by the paper.

Zhisheng-researcher/VeinCast implementation
Open GitHub

OWASP IoTGoat is deliberately insecure OpenWrt-based firmware containing vulnerability challenges mapped to the OWASP IoT Top 10.

OWASP/IoTGoat benchmark
Open GitHub

A repository supplying complementary medical research skills discussed in the capability survey and used as part of the paper's skill ecosystem.

aipoch/medical-research-skills
Open GitHub

The OpenClaw medical skills collection analyzed for its architecture, categories, composition, deployment model, and biomedical capability coverage.

FreedomIntelligence/OpenClaw-Medical-Skills
Open GitHub

Official API and supporting resources for running the VirtualHome household simulator and executing activity programs.

xavierpuigf/virtualhome dataset
Open GitHub

Unity source code for constructing VirtualHome environments and translating activity programs into low-level executable character actions.

xavierpuigf/virtualhome_unity
Open GitHub

Repository containing raw collected data, Study 1 to Study 3 folders, preregistration documents, analysis scripts written and executed by the system, analysis outputs, and manuscripts.

Explore-Science/Virtuous-Machines-Towards-Artificial-General-Science dataset
Open GitHub

Repository for the Matterport3D Simulator; it includes the R2R task assets/evaluation code and identifies this paper as the reference for the simulator and dataset.

peteanderson80/Matterport3DSimulator benchmark
Open GitHub

DMHouse, the modified DeepMind Lab and Quake-based indoor office simulator used to generate varied synthetic navigation environments for pre-training.

jkulhanek/dmlab-vn
Open GitHub

Official implementation of the proposed navigation method, including simulation and TurtleBot training and evaluation code, model configurations, and links to released checkpoints and the real-world dataset.

jkulhanek/robot-visual-navigation
Open GitHub

Official repository linked by the paper for Visual Semantic Entropy.

tadeephuy/visual-semantic-entropy
Open GitHub

Official repository for the VLQA benchmark, project webpage, dataset access, and associated code resources.

shailaja183/vlqa dataset
Open GitHub

Repository identified by the paper as providing the ViTOED source code and dataset.

sonlam1102/vitoed dataset
Open GitHub

Official code repository for VLIS, including inference interfaces for BLIP-2, LLaVA, and Lynx and data and evaluation code for the landmark and character benchmarks.

JiwanChung/vlis
Open GitHub

Repository for the VolleyBots environment, task suite, baseline implementations, and reproducible benchmark code.

thu-uav/VolleyBots benchmark
Open GitHub

Official implementation repository for the VoQA task, benchmark construction, evaluation, and question-alignment fine-tuning artifacts introduced by the paper.

AJN-AI/VoQA benchmark
Open GitHub

Official Voyager codebase, including the agent implementation and learned skill-library resources used to reproduce or apply the system described in the paper.

MineDojo/Voyager
Open GitHub

Official repository for the VQA^2 instruction dataset, model family, training and evaluation implementation, and released project artifacts.

Q-Future/Visual-Question-Answering-for-Video-Quality-Assessment
Open GitHub

Repository containing code associated with W-RAG weak-label generation and retriever fine-tuning experiments.

jmnian/WRAG framework
Open GitHub

Repository containing code and data for BadAgents, including poisoned data and code paths for Query-Attack, Observation-Attack, and Thought-Attack experiments.

lancopku/agent-backdoor-attacks dataset
Open GitHub

Continuously updated repository for tracking related works associated with the survey's human-view video-understanding taxonomy.

marinero4972/Awesome-HumanView-VideoUnderstanding
Open GitHub

Official OpenAI repository implementing a re-creation of the weak-to-strong binary-classification setup, supported losses including the confidence auxiliary loss, and the supplementary AlexNet-to-DINO vision experiment.

openai/weak-to-strong implementation
Open GitHub

The GitHub repository hosts the WebArena code, browser environment, evaluation harness, configuration files, Docker environment resources, prompts, scripts, and reproduction materials for the paper.

web-arena-x/webarena dataset
Open GitHub

Repository for the WebShop environment, product and instruction setup, search engine, baseline rule/IL/RL models, tests, and sim-to-real transfer code.

princeton-nlp/WebShop benchmark
Open GitHub

Repository containing MEDQA data and baseline source code associated with the paper.

jind11/MedQA dataset
Open GitHub

Repository for the ICLR 2025 AI feedback tool that evaluated submitted reviews for vagueness or genericity, possible misunderstanding of the paper, and unprofessional tone, then generated private improvement suggestions.

zou-group/review_feedback_agent system
Open GitHub

Official code repository for SCRL, including the verl-based implementation, preprocessing instructions, and a training launch example.

Jasper-Yan/SCRL implementation
Open GitHub

LLMAgora, the configurable arena used to run two-agent scenarios with public utterances, private reflections, surveys, parameter sweeps, logging, and optional semantic, NLI, and emotion analyses.

danmohad/LLMAgora
Open GitHub

Official repository containing code and data for the paper, including data preparation, reasoning-feature extraction, regression analysis, SAE analysis workflow, and test-time selection.

dayeonki/multilingual_reasoning dataset
Open GitHub

The robomimic repository contains the framework, algorithms, dataset interfaces, experiment configurations, and tooling associated with the paper's reproducible study.

ARISE-Initiative/robomimic benchmark
Open GitHub

Paper-linked GitHub URL for the BigData22 stock-movement dataset; the URL returned 404 during extraction, so repository availability was not verified.

stocktweet/stock-tweet dataset
Open GitHub

Repository associated with Hybrid Deep Sequential Modeling for Social Text-Driven Stock Prediction and its dataset.

wuhuizhe/CHRNN dataset
Open GitHub

Repository releasing a stock movement prediction dataset from tweets and historical stock prices.

yumoxu/stocknet-dataset dataset
Open GitHub

Repository for StockAgent, the LLM-based multi-agent stock trading simulation framework studied in the paper.

MingyuJ666/Stockagent framework
Open GitHub

A research-grade signal-only decision-support system for cross-sectional ranking of AI-focused U.S. equities with uncertainty quantification, regime-aware deployment gating, PIT-safe data handling, and walk-forward evaluation outputs.

sinsasanderink/AIStockForecaster-PIT-Safe-Ranking-First-Signals-for-AI-Equities-FMP-Kronos-FinText-TSFM- framework
Open GitHub

Repository containing code for reasoning-task experiments, computer-algebra calculations used in the appendices, and GPT-2/WikiText experiments associated with the paper.

eboix/relational-reasoning implementation
Open GitHub

Official implementation of OntoGraphRAG v1.0.0, the experiment harness, scripts for all tables and figures, per-query run logs, GPS replay stores, robustness artefacts, and reproducibility manifests used by the paper.

julka01/OntoGraphRAG
Open GitHub

MM-TSFlib implementation used for Informer, FEDformer, PatchTST, iTransformer, and DLinear backbone experiments.

AdityaLab/MM-TSFlib benchmark
Open GitHub

TimeCMA codebase used as an evaluated aligning-based MMTS comparison.

ChenxiLiu-HNU/TimeCMA implementation
Open GitHub

LeRet codebase used as an evaluated aligning-based MMTS comparison.

hqh0728/LeRet implementation
Open GitHub

Time-LLM codebase used as an evaluated aligning-based comparison model.

KimMeen/Time-LLM implementation
Open GitHub

Codebase for the Context is Key forecasting benchmark, reproduced to examine LLM performance scaling.

ServiceNow/context-is-key-forecasting benchmark
Open GitHub

Repository containing an anonymized CAIA evaluator implementation, a benchmark.csv dataset with 178 evaluation questions, evaluation scripts for with-tool and without-tool settings, mock tools, prompts, and dependencies.

caiba-ai/caia-benchmark-0927 dataset
Open GitHub

Car Crash Dataset (CCD), used by the paper to validate the vehicular accident-report use case based on mobile and edge LLM agents.

Cogito2012/CarCrashDataset dataset
Open GitHub

Repository containing classification and generation code for extracting linguistic and internal confidence, running ablations and analyses, preprocessing generation datasets, and reproducing reported confidence-correlation results.

HF-heaven/Correlation-between-Confidence-Measurements
Open GitHub

Repository containing the compact RLHF pipeline, transition classifier, experiment configurations, analysis scripts, manuscript sources, result tables, figures, examples, and a Gradio-based response-comparison interface.

zabahana/rlhf-failure-modes-diagnostics framework
Open GitHub

Repository identified as the code for 'When Routing Collapses: On the Degenerate Convergence of LLM Routers'.

AIGNLAI/EquiRouter implementation
Open GitHub

Repository containing frozen audit artifacts and code used to reproduce the reported NAVSIM score-basis conditions and numerical diagnostics.

WZiang/navsim-score-basis-audit
Open GitHub

Repository stated by the paper as the location where code and data for AgentDebug will be available.

ulab-uiuc/AgentDebug dataset
Open GitHub

NVIDIA's transformer inference engine, extended in the paper with DeepSpeed MoE support, TUPE attention, expert routing, quantized MoE computation, and batch pruning.

NVIDIA/FasterTransformer framework
Open GitHub

OpenManus repository cited by the paper for one of the multi-agent frameworks evaluated in MAST-Data.

mannaandpoem/OpenManus
Open GitHub

Official repository containing code and data for the paper, including taxonomy definitions/examples, traces, inter-annotator annotations, and the LLM judge pipeline.

multi-agent-systems-failure-taxonomy/MAST
Open GitHub

Repository containing prompts, research ideas, selected outputs, and failure analyses for the four autonomous research attempts, including MARL-idea, SALVO-WM-idea, SDTS-WM-idea, SemEnt-ALGN-idea, and workflow prompts.

Lossfunk/ai-scientist-artefacts-v1
Open GitHub

Author-linked repository for the WikiHow summarization dataset introduced by the paper.

mahnazkoupaee/WikiHow-Dataset dataset
Open GitHub

Repository for preparing, analyzing, and visualizing survey responses for the 'Will Agents Replace Us?' preprint project, including Python scripts for data preparation, exploratory analysis, inferential analysis, and manuscript figure generation.

nkkko/agent-perceptions dataset
Open GitHub

Repository for DeepFund, a platform intended to evaluate LLM trading capability across financial markets using a unified environment, multi-agent system, external information ingestion, trading decisions, and arena-style performance presentation.

HKUSTDial/DeepFund framework
Open GitHub

Repository for Latency Sensitive Benchmarks, including HFTBench and StreetFighter benchmark code and evaluation examples for latency-aware LLM-agent assessment.

HaoKang-Timmy/LatencySensitiveBench benchmark
Open GitHub

Repository linked by the paper for the proposed automated wireless-agent workflow design system.

jwentong/WirelessAgent-R2 framework
Open GitHub

Repository for WirelessBench, including the wireless tasks and released scoring code.

jwentong/WirelessBench dataset
Open GitHub

ParlAI repository containing the Wizard of Wikipedia task integration and project code used to distribute and evaluate the benchmark and associated models.

facebookresearch/ParlAI
Open GitHub

Official project repository containing the WizardCoder code/resources and linking the ICLR 2024 WizardCoder paper.

nlpxucan/WizardLM
Open GitHub

Framework used for running, managing, and reproducing web-agent experiments on BrowserGym benchmarks.

ServiceNow/AgentLab implementation
Open GitHub

Open-source Gym-style browser environment for implementing and evaluating web agents with rich observations and action spaces.

ServiceNow/BrowserGym framework
Open GitHub

Open-source benchmark package for evaluating browser agents on ServiceNow-based knowledge-work tasks.

ServiceNow/WorkArena dataset
Open GitHub

OpenBMB/WorkflowLLM is the official repository for the WorkflowLLM project, described as a data-centric framework for enhancing LLM workflow orchestration with WorkflowBench and WorkflowLlama resources.

OpenBMB/WorkflowLLM dataset
Open GitHub

Official repository for World Action Planner, including environments, world-model code/client, setup instructions, checkpoints, and a demo notebook for imagined actions.

XiangchengZhang/world-action-planner implementation
Open GitHub

Repository for the W2C pipeline, generated-data workflow, and associated LLaVA training and evaluation materials.

foundation-multimodal-models/World2Code
Open GitHub

Official implementation artifact for Worldscape-MoE, including training and inference entry points, modality-specific data preparation, validation tools, and optional offline VAE encoding.

EmbodiedCity/Worldscape-MoE.code
Open GitHub

Official code repository for training and evaluating Direct Preference Head models, including benchmark evaluation scripts and links to released model checkpoints.

Avelina9X/direct-preference-heads implementation
Open GitHub

NEORV32 is the architecturally distinct target SoC used to test whether training supervision generated from PicoSoC specifications transfers without target-design supervision.

stnolting/neorv32
Open GitHub

PicoSoC, a simple SoC built around PicoRV32, is the source architecture for the zero-shot IP-level transfer study.

YosysHQ/picorv32
Open GitHub

Repository for the paper's XGBoost-based NEPSE log-return forecasting workflow and benchmark outputs.

sahajrajmalla/nepse-xgboost-forecasting dataset
Open GitHub

Public repository containing the cost-aware LLM routing system, training/data-preprocessing components, evaluation and serving pipeline, router tests, and documentation.

SalesforceAIResearch/xRouter framework
Open GitHub

Code, scripts, notebooks, data assets, and FixMatch integration for reproducing DIPS tabular and computer-vision experiments.

seedatnabeel/DIPS implementation
Open GitHub

Official implementation of ZipCache, including package setup and an inference demonstration.

ThisisBillhe/ZipCache
Open GitHub

Official implementation of ZSMerge with model-specific attention modules, tests, datasets, and scripts for throughput and ROUGE experiments.

SusCom-Lab/ZSMerge
Open GitHub

Paper Details