LLM-Powered Agents - Index

Routing Summary

Large language models as the decision engine of simulated agents: generative agents with memory, “silicon samples” of survey respondents, Homo silicus economic subjects, and the validity problems all three share. Anchored by Gao et al. (2023 survey), Park et al. (2023), Argyle et al. (2022), Horton, Filippas & Manning (2023/2026) and Santurkar et al. (2023). Complements the rule-based consumer ABM, calibration, WOM and diffusion notes elsewhere in Agent-Based Modeling.

Concept Map

ConceptNoteTypeDepends OnKey Result
LLM-empowered ABM paradigmLLM-Powered Agents - OverviewoverviewABM Methodology; Agent Decision Rules; HeterogeneityFour LLM abilities (perception, reasoning, adaptation, heterogeneity) and four challenges (environment, alignment/personalisation, action simulation, evaluation); micro vs. macro evaluation
Generative agentsGenerative Agents Architecture - Memory, Reflection and PlanningmethodOverview; Emergent Phenomenascore = recency () + importance (LLM 1–10) + relevance (cosine); reflection at importance sum > 150; party awareness 4%→52%, density 0.167→0.74; full architecture TrueSkill 29.89 vs 21.21,
Silicon samplingSilicon Samples and Algorithmic FidelityconceptOverview; Poststratification; four fidelity criteria; tetrachoric 0.90/0.92/0.94; Cramér’s mean diff
Homo silicusHomo Silicus - LLMs as Simulated Economic AgentsconceptSilicon Samples; Decision Rules; Discrete ChoiceFive recapitulations; “theory in flexibly executable form”; confusion-matrix view; results need empirical confirmation
Persona mixture calibrationPersona Mixture Calibration of LLM AgentsmethodHomo Silicus; ABM Calibration Overview on the simplex; out-of-sample MSE 0.094 vs 0.182
LLM vs rule-based agentsLLM Agents vs Rule-Based Agents in ABMconceptOverview; Decision Rules; Calibration; ValidationCalibration burden becomes validation burden; “judge and jury” and Lucas-critique arguments; hybrid designs
Opinion alignment metricsOpinion Alignment Metrics for Language ModelsmethodSilicon Samples; every human group beats every LM on representativeness; steering modest; RLHF modal collapse
Validity and calibrationValidity, Bias and Calibration of LLM-Simulated Populationsconceptall of the above; ABM Validation Challenges; Poststratification15-threat taxonomy; imputed-context confounding; 10-step validation protocol; prediction-powered inference

Notes

  • LLM-Powered Agents - Overview — CONTAINS: LLM-agent definition, classical agent desiderata, Gao et al.’s four abilities and four challenges, micro/macro evaluation, application domains, open problems (scaling, benchmarks, robustness, ethics), relevance to marketing measurement, minimal agent-step sketch.
  • Generative Agents Architecture - Memory, Reflection and Planning — CONTAINS: memory stream definition, retrieval scoring formula and constants, reflection algorithm and reflection trees, recursive planning and reaction, environment grounding, ablation results table, emergent diffusion/network/coordination numbers, failure modes, retrieval code sketch.
  • Silicon Samples and Algorithmic Fidelity — CONTAINS: LM-as-conditional-distribution, algorithmic fidelity definition, four criteria, silicon-sampling marginal-correction identity and its equivalence to poststratification, Studies 1–3 results, critical reading, worked reweighting example, pipeline sketch.
  • Homo Silicus - LLMs as Simulated Economic Agents — CONTAINS: Homo silicus definition, latent vs. explicit social information, five recapitulated experiments with designs and numbers, Charness–Rabin utility, theory-as-instruction argument, Fig. 7 confusion matrix, prediction-powered inference, when simulations are informative, in-silico pilot sketch.
  • Persona Mixture Calibration of LLM Agents — CONTAINS: theory-grounded persona definition, six-step mixture algorithm, reported weights per model, out-of-sample MSE, comparison table of calibration strategies, identification/uncertainty caveats with Bayesian extension, worked constrained-least-squares example.
  • LLM Agents vs Rule-Based Agents in ABM — CONTAINS: side-by-side comparison table, Gao’s four advantages, Horton’s “judge and jury” and Lucas-critique arguments, cost catalogue, four hybrid design patterns, same-decision two-agent example.
  • Opinion Alignment Metrics for Language Models — CONTAINS: OpinionQA construction, human/model opinion distributions, Wasserstein alignment (Eq. 1), representativeness/steerability/consistency definitions, empirical findings, implications for ABM, hand-computed alignment examples, metric code.
  • Validity, Bias and Calibration of LLM-Simulated Populations — CONTAINS: 15-row threat taxonomy with source evidence, poststratification limits, social-influence bias, imputed-context confounding in potential-outcomes terms, 10-step protocol, four-outcomes definition, disagreements across sources, brand-tracker checklist, dispersion check code.

Sources

  • Gao 2023 - LLM Empowered Agent-Based Modeling Survey — Gao, C., Lan, X., Li, N., Yuan, Y., Ding, J., Zhou, Z., Xu, F. & Li, Y. (2023), “Large Language Models Empowered Agent-based Modeling and Simulation: A Survey and Perspectives,” arXiv:2312.11970.
  • Park 2023 - Generative Agents Interactive Simulacra — Park, J. S., O’Brien, J. C., Cai, C. J., Morris, M. R., Liang, P. & Bernstein, M. S. (2023), “Generative Agents: Interactive Simulacra of Human Behavior,” UIST ‘23, arXiv:2304.03442.
  • Argyle 2022 - Out of One Many Silicon Samples — Argyle, L. P., Busby, E. C., Fulda, N., Gubler, J., Rytting, C. & Wingate, D. (2022), “Out of One, Many: Using Language Models to Simulate Human Samples,” arXiv:2209.06899.
  • Horton 2023 - Homo Silicus LLMs as Simulated Economic Agents — Horton, J. J., Filippas, A. & Manning, B. S. (2023; v2 2026), “Large Language Models as Simulated Economic Agents: What Can We Learn from Homo Silicus?”, arXiv:2301.07543.
  • Santurkar 2023 - Whose Opinions Do Language Models Reflect — Santurkar, S., Durmus, E., Ladhak, F., Lee, C., Liang, P. & Hashimoto, T. (2023), “Whose Opinions Do Language Models Reflect?”, arXiv:2303.17548.