When Images Pretend to Be Interfaces: Stress‑Testing Generative Models as GUI Environments
GEBench shows why beautiful generated interfaces are not yet reliable environments for training or testing GUI agents.
GEBench shows why beautiful generated interfaces are not yet reliable environments for training or testing GUI agents.
A careful reading of FedCompDP shows why privacy, client heterogeneity, and aggregation stability must be designed together—not bolted together after the model starts shaking.
CompactRAG shows how multi-hop RAG can shift cost from repeated online LLM calls to reusable offline knowledge compaction.
TimelyFreeze shows that parameter freezing only becomes a real training-speed lever when it is aligned with the pipeline schedule’s wall-clock bottlenecks.
AutoInject shows why prompt injection should be tested as an adaptive optimization problem, not merely as a list of hand-written attack templates.
VSD shows why speculative decoding improves when draft models are trained for accepted paths, not merely probable tokens.
A mechanism-first reading of why LLM inference energy is shaped by prefill, decoding, prompt length, and unnecessary generation—not merely model size.
CSRv2 shows that ultra-sparse embeddings fail less because sparsity is impossible, and more because we have been training them badly.
A comparison-based reading of why Word Mover’s Distance with GloVe outperforms centroid-style semantic search in statement-level retrieval, and where that lesson actually applies in business systems.
A mechanism-first reading of TEA, an in-situ task-generation framework showing why embodied AI needs environment-specific evaluation before deployment.