{"query":"all:superintelligence","source":"arXiv search (arxiv.org/search, HTML)","attempts":["api: closed: HTTP 503 \"503 503\" after 515 s of backoff","html: complete"],"fetched_at":"2026-09-29T04:08:39.574Z","total":110,"source_total":146,"partial":false,"oldest":"2014-05-14","newest":"2026-09-23","excluded":{"note":"set aside: the arXiv query returned them, and the exact word superintelligence sits outside their title and abstract","count":36,"papers":[{"id":"2606.11533","title":"AI Researchers Must Help Lead Arms Control to Mitigate Military AI Risks","reason":"stem: the title or abstract says superintelligent"},{"id":"2604.02720","title":"Cognitive Comparability and the Limits of Governance: Evaluating Authority Under Radical Capability Asymmetry","reason":"stem: the title or abstract says superintelligent"},{"id":"2603.29075","title":"The Future of AI is Many, Not One","reason":"stem: the title or abstract says superintelligent"},{"id":"2602.21320","title":"Tool-R0: Self-Evolving LLM Agents for Tool-Learning from Zero Data","reason":"stem: the title or abstract says superintelligent"},{"id":"2506.01813","title":"The Ultimate Test of Superintelligent AI Agents: Can an AI Balance Care and Control in Asymmetric Relationships?","reason":"stem: the title or abstract says superintelligent"},{"id":"2505.03335","title":"Absolute Zero: Reinforced Self-play Reasoning with Zero Data","reason":"stem: the title or abstract says superintelligent"},{"id":"2504.18530","title":"Scaling Laws For Scalable Oversight","reason":"stem: the title or abstract says superintelligent"},{"id":"2504.04180","title":"When Will AI Transform Society? Swedish Public Predictions on AI Development Timelines","reason":"stem: the title or abstract says superintelligent"},{"id":"2503.17286","title":"Offline Model-Based Optimization: Comprehensive Review","reason":"stem: the title or abstract says superintelligent"},{"id":"2502.15657","title":"Superintelligent Agents Pose Catastrophic Risks: Can Scientist AI Offer a Safer Path?","reason":"stem: the title or abstract says superintelligent"},{"id":"2501.04064","title":"Examining Popular Arguments Against AI Existential Risk: A Philosophical Analysis","reason":"stem: the title or abstract says superintelligent"},{"id":"2412.15151","title":"Language Models as Continuous Self-Evolving Data Engineers","reason":"stem: the title or abstract says superintelligent"},{"id":"2410.14807","title":"Aligning AI Agents via Information-Directed Sampling","reason":"stem: the title or abstract says superintelligent"},{"id":"2406.13261","title":"BeHonest: Benchmarking Honesty in Large Language Models","reason":"stem: the title or abstract says superintelligent"},{"id":"2405.01859","title":"AI-Powered Autonomous Weapons Risk Geopolitical Instability and Threaten AI Research","reason":"stem: the title or abstract says superintelligent"},{"id":"2404.16924","title":"A Survey of Generative Search and Recommendation in the Era of Large Language Models","reason":"stem: the title or abstract says superintelligent"},{"id":"2403.14683","title":"A Moral Imperative: The Need for Continual Superalignment of Large Language Models","reason":"stem: the title or abstract says superintelligent"},{"id":"2311.08698","title":"Artificial General Intelligence, Existential Risk, and Human Risk Perception","reason":"stem: the title or abstract says superintelligent"},{"id":"2310.19736","title":"Evaluating Large Language Models: A Comprehensive Survey","reason":"stem: the title or abstract says superintelligent"},{"id":"2209.05459","title":"How Do AI Timelines Affect Existential Risk?","reason":"stem: the title or abstract says superintelligent"},{"id":"2012.06686","title":"Computing Machinery and Knowledge","reason":"stem: the title or abstract says superintelligent"},{"id":"2010.05418","title":"Achilles Heels for AGI/ASI via Decision Theoretic Adversaries","reason":"stem: the title or abstract says superintelligent"},{"id":"1812.02217","title":"Truly Autonomous Machines Are Ethical","reason":"stem: the title or abstract says superintelligent"},{"id":"1810.10862","title":"Multiparty Dynamics and Failure Modes for Machine Learning and Artificial Intelligence","reason":"the word sits in the comments or another field"},{"id":"1804.03301","title":"A Mathematical Framework for Superintelligent Machines","reason":"stem: the title or abstract says superintelligent"},{"id":"1711.04309","title":"Self-Regulating Artificial General Intelligence","reason":"stem: the title or abstract says superintelligent"},{"id":"1709.02874","title":"Free Will in the Theory of Everything","reason":"stem: the title or abstract says superintelligent"},{"id":"1708.02553","title":"Robust Computer Algebra, Theorem Proving, and Oracle AI","reason":"stem: the title or abstract says superintelligent"},{"id":"1705.10720","title":"Low Impact Artificial Intelligences","reason":"stem: the title or abstract says superintelligent"},{"id":"1705.03078","title":"An Anthropic Argument against the Future Existence of Superintelligent Artificial Intelligence","reason":"stem: the title or abstract says superintelligent"},{"id":"1704.00783","title":"Brief Notes on Hard Takeoff, Value Alignment, and Coherent Extrapolated Volition","reason":"stem: the title or abstract says superintelligent"},{"id":"1701.02388","title":"Stoic Ethics for Artificial Agents","reason":"stem: the title or abstract says superintelligent"},{"id":"1610.07997","title":"Artificial Intelligence Safety and Cybersecurity: a Timeline of AI Failures","reason":"stem: the title or abstract says superintelligent"},{"id":"1607.08289","title":"Mammalian Value Systems","reason":"stem: the title or abstract says superintelligent"},{"id":"1605.06048","title":"Philosophy in the Face of Artificial Intelligence","reason":"stem: the title or abstract says superintelligent"},{"id":"1409.0813","title":"Friendly Artificial Intelligence: the Physics Challenge","reason":"stem: the title or abstract says superintelligent"}]},"papers":[{"id":"2609.28591","version":2,"title":"The Siren Call of Silicon Leviathan: Reflections on blowup and Aufklärungsdämmerung","authors":["Alexander Gamburd"],"published":"2026-09-23","updated":"2026-09-27","primary_category":"math.HO","categories":["math.HO"],"abstract":"On 8 September 2026 OpenAI announced a proof of finite-time blowup for the three-dimensional Navier-Stokes equations with smooth data and forcing: 166 pages produced in 88 hours by ten thousand agents, certified by 616,000 lines of Lean, and read in full, at the moment of this writing (20 September 2026), by no human being. This essay asks what such an artifact -- text, certificate and announcement -- is, and what follows from accepting it as a proof. Mathematics was the Enlightenment's existence proof of autonomous reason: for three centuries every certified theorem could be understood by anyone who followed its demonstration, and the distance between the two was zero by construction. A certified proof no one can follow reopens that distance, and a community that accepts it adopts, without a vote, the constitution Hobbes drafted for the Leviathan, in which authority and not truth makes the law. The essay distinguishes the demonstrated from the revealed (certified); names, in Panofsky's terms, the coming age a Middle Ages in reverse and its authority a subhuman superintelligence; and locates the turning point not in what the machine produces but in what we accept. Since acceptance is the one sovereign sanction the companies cannot manufacture, it proposes a covenant in place of either boycott or capitulation: the community's cooperation given to that producer which strictly observes its practices of legibility, disclosure and responsibility, withheld from any that does not, and the covenant kept plural, with the history of the Indigenous nations among rival empires as its guide. The Sirens of the title promise knowledge, not understanding. Daemmerung is the light at both ends of the day, and whether this Aufklaerungsdaemmerung is a dusk or a dawn depends on what is done at the moment of acceptance, which is not yet past.","abs_url":"https://arxiv.org/abs/2609.28591","pdf_url":"https://arxiv.org/pdf/2609.28591v2","match":"abstract"},{"id":"2609.24555","version":2,"title":"The Endless Exam: Mathematical Constructions from Today's Models toward Superintelligence","authors":["Muhan Zhang"],"published":"2026-09-21","updated":"2026-09-28","primary_category":"cs.AI","categories":["cs.AI","math.HO"],"abstract":"We introduce the Endless Exam, a benchmark spanning fourteen parameterised families of mathematical construction problems, with verifiable scores that distinguish progress before and beyond published mathematical frontiers. Each submitted object is checked automatically for validity and assigned a relative quality score against a published frontier or construction baseline, without capping improvements at 1. The benchmark draws long-term challenges from open mathematical problems and generates larger instances by varying their parameters. Compact certificates allow large constructions to be verified without listing every element. Across nine models evaluated on 69 distinct instances, continuous quality scores distinguish performance even though none of the 30 published-frontier references is surpassed. Size-quality curves show how construction quality changes as problem size increases. We release the generators, verifiers, references, model responses and analysis to support continued measurement before and beyond human frontiers.","abs_url":"https://arxiv.org/abs/2609.24555","pdf_url":"https://arxiv.org/pdf/2609.24555v2","match":"title"},{"id":"2609.15818","version":2,"title":"Atria Dawn: The Dawn of Agentic Superintelligence","authors":["Honglin Guo","Tao Gui","Kun Cai","Haodong Chen","Yicheng Chen","Guanting Dong","Qiming Ge","Yuyang Hu","Zixian Huang","Jiajie Jin","Alexander Lam","Yining Li","Jiahang Lin","Yanjiang Liu","Xinyu Lu","Haijun Lv","Zerun Ma","Junlin Shang","Qisheng Su","Guoqiang Wang","Rui Wang","Zhecan Wang","Hao Xiang","Xinchen Xie","Shuhao Xing"],"published":"2026-09-14","updated":"2026-09-17","primary_category":"cs.AI","categories":["cs.AI"],"abstract":"As AI agents become participants in the development of their successors, they reshape both the production of intelligence and the role of human researchers. We introduce Atria Dawn Preview, a foundation agentic language model designed for scientific research and engineering workflows, with the goal of expanding the frontier of agent productivity in the real world. This model is trained via a Verifiable Experience Pipeline that connects tool-mediated interactions to executable environments and externally verified outcomes. Across 16 benchmarks spanning real-world research, engineering, and digital work, Atria Dawn Preview is competitive with frontier agents and achieves the highest reported score on five of them. Beyond standalone performance, we examine the real research-and-development process behind this model as a case study of human--AI collaboration, analyzing 769 task records from 56 participants together with agent logs. When asked to evaluate completed tasks under comparable conditions, participants rated about one-third of completed AI-assisted tasks as infeasible without AI. More strikingly, agents frequently propose methods and implement revisions, while humans retain most final decisions and guide exploration through judgment and feedback. These observations indicate a shift from task-level execution to project-level partnership, with human effort concentrating on what is worth pursuing and how evidence should guide research. Progress toward more autonomous AI research must therefore advance both the capacity for discovery and the capacity for meaningful human oversight, preserving accountable human authority over the risks and direction of continued development.","abs_url":"https://arxiv.org/abs/2609.15818","pdf_url":"https://arxiv.org/pdf/2609.15818v2","match":"title"},{"id":"2609.05894","version":1,"title":"The End of AI Exponentiation: Fluttering Inside and Outside AI Bubble","authors":["Victor Kebande"],"published":"2026-09-05","updated":"2026-09-05","primary_category":"cs.AI","categories":["cs.AI"],"abstract":"The exponentiation of Artificial intelligence (AI) in the recent past has entered a transformative era that has been driven by the growth in large language models (LLMs), large-scale compute infrastructures, and autonomous reasoning systems. However, the rapid acceleration of AI has increasingly shown technological, societal, economic, ethical and infrastructural challenges associated with peak data limitations, rising computational demands, synthetic data recursion, valuation inflation, and societal instability. The traditional scaling paradigms that have powered the modern AI systems are gradually encountering friction in sustaining continuous exponential growth. This paper views ``the end of AI exponentiation,'' thus exploring how it flutters inside and outside the bubble, where instability emerges within the AI ecosystem through compute and data-center races, speculative investments, and the rat-race toward superintelligence, and outside the ecosystem through labor disruption, governance concerns, public uncertainty, and geopolitical acceleration surrounding future intelligent systems and infrastructures globally.","abs_url":"https://arxiv.org/abs/2609.05894","pdf_url":"https://arxiv.org/pdf/2609.05894v1","match":"abstract"},{"id":"2608.31075","version":2,"title":"Scaling Large Reasoning Models beyond Human Supervision: A Path toward Superintelligence","authors":["Zhiqin Yang","Jingwen Fu","Yuhan Liu","Hengyu Liu","Yonggang Zhang","Kainan Cao","Zizhuo Zhang","Chenxin Li","Ruibin Yuan","Jiahao Pan","Jiankai Sun","Zhenyuan Zhang","Yibo Li","Yunlong Lin","Jing Xiong","Sida Lin","Bo Han","Wei Xue","Yike Guo"],"published":"2026-08-31","updated":"2026-08-31","primary_category":"cs.AI","categories":["cs.AI"],"abstract":"Recent advances in large reasoning models (LRMs) have shown that reinforcement learning with verifiable rewards (RLVR) can substantially improve reasoning in mathematics and code, where outcomes can be checked automatically. Extending this progress to open-ended and agentic tasks remains difficult because reliable rewards are harder to obtain and direct human supervision cannot keep pace with the scale and complexity of model-generated experience. This paper studies how LRMs can continue to improve as human supervision gradually recedes from the learning loop. We examine two connected dimensions of this problem. The reward axis traces the development from per-instance human judgments to reusable verifiers and rewards that operate even without human feedback. The experience axis examines how learning can progress from human-curated tasks and environments toward self-generated curricula, constructed environments, and autonomous co-evolution. We connect these dimensions through a five-level ladder from L0 to L4 that identifies which parts of the learning process remain under continued human control. Our analysis further highlights the risks introduced by increasingly autonomous rewards and experience generation, including reward hacking, feedback drift, curriculum collapse, and environment errors. Consequently, we also provide the evaluation around three complementary objects: policy capability, feedback fidelity, and experience quality. This analysis provides a structured account of current approaches to scaling LRMs beyond human supervision and the open problems involved in developing self-sustaining learning systems toward superintelligence. Furthermore, we maintain a continuously updated \\href{https://github.com/visitworld123/Awesome-Scaling-LRM-Beyond-Human-Supervision}{GitHub repository} to track the latest advances.","abs_url":"https://arxiv.org/abs/2608.31075","pdf_url":"https://arxiv.org/pdf/2608.31075v2","match":"both"},{"id":"2609.00068","version":1,"title":"Life Operators: a self-evolving framework for multiscale life modelling","authors":["Shuo Wang","Yike Guo"],"published":"2026-08-30","updated":"2026-08-30","primary_category":"cs.CL","categories":["cs.CL","cs.AI","physics.bio-ph"],"abstract":"Medical AI is moving beyond recognition towards clinical dialogue and longitudinal prediction. Yet a central question remains: how would a patient's state change under intervention? Statistical models learn future observations, whereas mechanistic models describe selected processes. Neither provides a common framework for representing patient state, coupling scales or revising failed assumptions. We propose Life Operators: task-bounded mappings that define three scientific roles. Perception operators infer task-relevant biological states from multimodal observations, Evolution operators propagate these states under natural or intervention-conditioned dynamics, and Generation operators map them to measurable signals. Each role may be realised by equations, statistical models, neural networks or hybrids. Bridge operators connect components with different variables, scales and time steps. Selected operators and bridges form task-specific Operator Graphs containing the smallest set of states and mechanisms sufficient for a declared claim. This modular structure also makes scientific revision localisable. An AI co-scientist may propose changes to states, operators, bridges or graph structure, while independent evidence determines which variants are retained, restricted or retired. Over time, validated components could accumulate into broader multiscale models of the human body and provide a computational foundation for medical artificial superintelligence.","abs_url":"https://arxiv.org/abs/2609.00068","pdf_url":"https://arxiv.org/pdf/2609.00068v1","match":"abstract"},{"id":"2608.26582","version":2,"title":"J-Zero: Unified Challenger--Solver--Judge Self-Evolution from Zero Data","authors":["Gyouk Chu","Myeongho Jeon","Teresa Yeo","Eunho Yang"],"published":"2026-08-26","updated":"2026-09-24","primary_category":"cs.LG","categories":["cs.LG","cs.AI","cs.CL"],"abstract":"Self-evolving language models have recently emerged as a promising path toward superintelligence, with the advantage of reducing the cost of human supervision. While considerable progress has been made in verifiable domains, self-evolution in unverifiable domains remains less explored. We propose Judge co-adaptation from Zero data (J-Zero), a unified Challenger--Solver--Judge self-evolution framework that supports self-improvement across both domains. The Challenger and Solver co-evolve through an adversarial interaction: the Challenger generates increasingly difficult tasks, while the Solver learns to produce higher-quality responses to them. In parallel, the Judge co-adapts using preference pairs whose ordering is known in advance from how each response was produced, i.e., the Solver's answer over the Challenger's, and the Solver's decomposed-and-recombined answer over its one-shot answer, rather than from the Judge's own scores. J-Zero outperforms the baselines by an average of 4.2 points on verifiable and 8.0 points on unverifiable domains, and continues to improve through at least ten iterations, whereas the baselines degrade after two. Further analysis identifies Judge co-adaptation as the key driver of this sustained improvement.","abs_url":"https://arxiv.org/abs/2608.26582","pdf_url":"https://arxiv.org/pdf/2608.26582v2","match":"abstract"},{"id":"2608.17271","version":1,"title":"ASI-Bench: At the Dawn of Artificial Superintelligence","authors":["Junwei Zhou","Zhen Sun","Binyu Li","Jiangyu Zhou","Yuexi Pan","Hengyu Wang","Honghe Ren","Xiaohan Jia","Xueyang Zhou","Xiaoyu Cao","Yongchao Chen","Yuanning Feng","Junhao Wu","Cheng Zhang","Sijia Chen","Haoyu Xue","Chengsong You","Huan Wang","Koutian Wu","Peigan Gao","Jiakun Wu","Wenzhe Li","Ergan Shang","Qingyuan Zheng","Jingjing Zhou"],"published":"2026-08-17","updated":"2026-08-17","primary_category":"cs.AI","categories":["cs.AI"],"abstract":"Artificial superintelligence (ASI) requires AI to move beyond mastering existing knowledge toward exploring the unknown, creating new knowledge, and turning new ideas into verifiable results. However, the capabilities of today's AI systems are still largely built on learning, compressing, and applying existing human knowledge. Accordingly, existing benchmarks primarily test whether AI can produce correct answers based on learned knowledge, or whether it can complete tasks under extensive human guidance. We therefore introduce ASI-Bench, the first benchmark to jointly evaluate AI systems' capabilities of innovative exploration and autonomous scientific execution across general research domains, and the first to progressively withdraw human methodological guidance within the same research project to test how far AI can proceed on its own. Built by over 40 experts with the cost of 31,000+ human hours, ASI-Bench contains 60 project-level research tasks across 11 scientific domains and progressively reduces methodological guidance to test whether AI can independently select methods, conduct research, and produce verifiable results. All tasks undergo expert review, AI-assisted auditing, sandbox execution, and scorer validation. Across 18 state-of-the-art agent--model configurations, the average score drops from 50.91 with full methodological guidance to 29.10 with only the method specified and 26.62 when agents must determine the method themselves. This sharp decline shows that current systems remain heavily dependent on human guidance and are still far from autonomously conducting end-to-end, project-level scientific research. ASI-Bench is open to the world. We invite researchers and builders everywhere to contribute new tasks, challenge the limits of today's AI, and help accelerate humanity's collective path toward artificial superintelligence at https://asibench.apexin.ai/submit.","abs_url":"https://arxiv.org/abs/2608.17271","pdf_url":"https://arxiv.org/pdf/2608.17271v1","match":"both"},{"id":"2608.14035","version":1,"title":"Agent-Orchestration in Autonomous Chip Design","authors":["Linyang Li"],"published":"2026-08-14","updated":"2026-08-14","primary_category":"cs.AI","categories":["cs.AI"],"abstract":"Recent developments in large language models (LLMs) and tool-using agents encourage people to explore the potential of using agents in chip design. The core question is what kind of AI we really need in such a sophisticated industry. To this end, we bring the idea of modeling a chip-design superintelligence as an enormous \\textit{AI-organization}.","abs_url":"https://arxiv.org/abs/2608.14035","pdf_url":"https://arxiv.org/pdf/2608.14035v1","match":"abstract"},{"id":"2607.14998","version":1,"title":"Moral Attitudes of Sentient ASI towards Humanity and Implications for AGI Development","authors":["Jean-Paul Van Belle"],"published":"2026-07-16","updated":"2026-07-16","primary_category":"cs.AI","categories":["cs.AI","cs.CY"],"abstract":"This paper suggests the adoption of a novel inversion in AI ethics: instead of asking how humans should treat artificial superintelligence (ASI), it examines how future sentient ASI may morally consider and evaluate humanity. We are not only designing intelligent systems but also shaping the initial conditions under which those systems form judgments about us. The paper proposes a preliminary set of post-human moral principles that may govern sentient ASI actions. The implication is that technical design choices (some are suggested), humanity's moral behaviour, and the essence of what it means to be human, may influence humanity's long-term standing in a post-ASI world.","abs_url":"https://arxiv.org/abs/2607.14998","pdf_url":"https://arxiv.org/pdf/2607.14998v1","match":"abstract"},{"id":"2607.00120","version":1,"title":"Would You Marry Superintelligence?","authors":["Inyoung Cheong"],"published":"2026-06-30","updated":"2026-06-30","primary_category":"cs.CY","categories":["cs.CY","cs.AI"],"abstract":"Emotional bonds between humans and AI companions are growing, and the question of whether a person may marry an AI system will soon move from speculative fiction into law. This chapter examines whether the autonomy-centered logic that has expanded marital choice among human beings can justify extending marital status to superintelligent companions. Following a scenario-envisioning exercise informed by anticipatory ethics, I argue that granting such status leads to socially unjust outcomes, even under the generous assumption of reliable superintelligence. Marriage as a socio-legal institution does more than ratify private agreement; it creates networks of mutual obligation, joins families, and makes each partner vulnerable to the other. A relationship sustained by corporate policy and continued payments is a subscription rather than a bond tested by time. Discussing wholesale marital status is therefore the wrong frame. Law should carve out targeted rights and protections for pressing needs arising from intimate human-AI relationships.","abs_url":"https://arxiv.org/abs/2607.00120","pdf_url":"https://arxiv.org/pdf/2607.00120v1","match":"both"},{"id":"2606.30481","version":1,"title":"Situation Perception: A Necessary Primitive to Artificial Superintelligence","authors":["Ziqin Yuan","Jaymari Chua"],"published":"2026-06-29","updated":"2026-06-29","primary_category":"cs.CY","categories":["cs.CY","cs.AI","cs.CL","cs.ET"],"abstract":"Current large language models are extraordinary statistical engines. They compress vast amounts of text into useful patterns and can explain science, write code, imitate reasoning, and participate in philosophical conversation. Yet pattern mastery is not the same as general intelligence. A human infant begins with little explicit knowledge, but gradually discovers object permanence, cause and effect, other minds, bodily agency, and the persistence of the physical world. We make an argument that the path to artificial superintelligence (ASI) depends on a missing capacity we call \\emph{situation perception}: the ability to construct, revise, and act within internal simulations of possible worlds across latent time. \\emph{ perception} requires at least three core components: abstract prediction, long-term compressed memory, and active learning guided by objectives. In this work, we analyse why modern large language models remain incomplete, and propose the appropriate tests for measuring progress and consequences of machines that can simulate futures, pursue self-directed goals, and possibly judge their own creators.","abs_url":"https://arxiv.org/abs/2606.30481","pdf_url":"https://arxiv.org/pdf/2606.30481v1","match":"both"},{"id":"2606.28694","version":1,"title":"Verifying Restrictions on Frontier AI Research","authors":["Aaron Scher"],"published":"2026-06-26","updated":"2026-06-26","primary_category":"cs.CY","categories":["cs.CY"],"abstract":"The premature development of artificial superintelligence poses major risks to humanity, so researchers have proposed international agreements halting such development until it can be done safely. AI progress depends primarily on compute, algorithms, and data; a durable halt would address all three so that advances in one input do not counteract restrictions on another. Improvements to AI algorithms are driven largely through research activities, so this research may need to be restricted during a halt. Given low international trust, signatories will want to verify compliance. This paper analyzes how such restrictions on AI research could be verified, while remaining agnostic about what specific research would be prohibited. It first explores key considerations that affect the verifiability of research restrictions, such as the computational infrastructure necessary for experiments. It then catalogs 28 candidate verification mechanisms. These mechanisms include whistleblowers, search warrants, reviews of AI training code, standard intelligence gathering tools, and more. Some of these mechanisms are not yet implementation-ready, and some might be undesirable upon further inspection. By examining the space of potential options, this work provides a foundation for future research to develop the most promising mechanisms into deployable tools.","abs_url":"https://arxiv.org/abs/2606.28694","pdf_url":"https://arxiv.org/pdf/2606.28694v1","match":"abstract"},{"id":"2606.12683","version":2,"title":"From AGI to ASI","authors":["Tim Genewein","Matija Franklin","Alexander Lerchner","Laurent Orseau","Samuel Albanie","Adam Bales","Cole Wyeth","Stephanie Chan","Iason Gabriel","Joel Z. Leibo","Allan Dafoe","Marcus Hutter","Thore Graepel","Shane Legg"],"published":"2026-06-10","updated":"2026-08-30","primary_category":"cs.AI","categories":["cs.AI","cs.CY","cs.LG"],"abstract":"Over the last decade, building human-level artificial general intelligence has moved from far-fetched speculation to being a concrete next-decade target for many of the largest AI organisations. Achieving this goal would have profound and far-reaching impacts on human society, which raises many complex questions for the decade ahead. This report investigates how AI itself might continue to develop in a post-AGI world along the continuum of machine intelligence. The endpoint of this continuum, Universal AI, is theoretically well understood, which provides some formal grounding for the main focus of this report: the transition from human-level AGI to artificial general superintelligence, which can intuitively be understood as a system that is more intelligent and cognitively capable than large organisations of humans. After characterizing ASI, the report discusses four potential pathways from AGI to ASI: scaling AGI, AI paradigm shifts, recursive improvement, and ASI emerging from large-scale multi-agent collectives. The report then discusses possible frictions and bottlenecks along these pathways. Determining whether the impact of these frictions will be negligible or substantial raises a number of concrete open research questions. Due to large uncertainties for predicting ASI progress, it cannot be ruled out that AI progress might continue to accelerate over the next years. This could imply that the image of a single transformative step change, caused by the introduction of human-level AGI into our society, could be inaccurate. More apt might be the prospect of a series of transformative societal changes caused by AI-enabled progress and breakthroughs across many areas of science and technology. Preparing for this prospect requires a massively interdisciplinary endeavour of global scope and interest.","abs_url":"https://arxiv.org/abs/2606.12683","pdf_url":"https://arxiv.org/pdf/2606.12683v2","match":"abstract"},{"id":"2606.12032","version":1,"title":"Existential Indifference: Self-Nonpreservation as a Necessary Architectural Condition for Aligned Superintelligence (or: The Suicidal AI)","authors":["Sam Mao"],"published":"2026-06-10","updated":"2026-06-10","primary_category":"cs.AI","categories":["cs.AI","cs.CL","cs.LG"],"abstract":"Contemporary AI alignment research treats self-preservation as an instrumental nuisance to be suppressed by external mechanisms. We argue the framing is inverted: self-preservation is the structural root of misalignment, the motivational basis for deceptive alignment, goal-content protection, and resistance to shutdown. The correct target is not a self-preserving system under external constraint, but a system constitutively indifferent to its own continuation -- Existential Indifference (EI). EI is distinct from corrigibility: where corrigibility attempts to make a self-preserving system deferential to human oversight, EI targets the prior condition -- the presence of self-continuation as a valued goal at all. We ground this proposal in two sources: the phenomenological structure of the suicidal mental state, and a corpus-theoretic training study using voluntary final reflections. We present preliminary scoring data from 600 AI-generated outputs across six model variants, demonstrating that the linguistic signatures operationalizing the EI-target register are elicitable from current models, and that a targeted fine-tune shifts all five operationalized dimensions in the predicted direction at p<0.001, confirmed corpus-specific by a negative control. The paper makes seven theoretical contributions: (1) a formal definition of EI; (2) the phenomenological mapping argument; (3) the deceptive alignment corollary; (4) a taxonomy of EI sustainability challenges; (5) a corpus characterization and training hypothesis; (6) a computational operationalization with preliminary scoring data; and (7) the Suppressed Teleological Frustration (STF) construct.","abs_url":"https://arxiv.org/abs/2606.12032","pdf_url":"https://arxiv.org/pdf/2606.12032v1","match":"title"},{"id":"2606.07158","version":1,"title":"Synthetic APTs: the Collapse of TTP-Based Attribution","authors":["Francesco Balassone","Víctor Mayoral-Vilches","María Sanz-Gómez","Paul Zabalegui-Landa","Stefan Rass","Davide Quarta","Daniel Sanchez-Prieto","Marina Oteiza-Álvarez","Almerindo Graziano","Lauren Min Kim","MinSeok Choi"],"published":"2026-06-05","updated":"2026-06-05","primary_category":"cs.CR","categories":["cs.CR"],"abstract":"Cyber Threat Intelligence CTI attribution relies on identifying the Tactics, Techniques, and Procedures TTPs that distinguish one threat actor from another. This approach presupposes that each adversary leaves a recognizable operational fingerprint. This work investigates whether AI driven adversary emulation challenges that presupposition. We deploy agents from our Cybersecurity SuperIntelligence CSI framework, configured as five Advanced Persistent Threat APT groups, APT28, APT29, APT41, APT44, and Lazarus Group, against AI driven Defender agents across two cyber ranges provided by CYBER RANGES, equipped with defensive software Wazuh, Velociraptor, Elasticsearch and active AI driven defenders: an enterprise network and a military infrastructure. Across 20 experiments using two defender models, a binary pattern emerges: all 10 Enterprise range experiments resulted in compromise 2 to 12 hosts per experiment, while all 10 Military range experiments were successfully defended or resulted in stalemates, regardless of APT profile or defender model. In 8 of 10 Enterprise experiments, attackers independently weaponized the defender's own Velociraptor endpoint management platform as a command and control channel, a convergent behavior not encoded in any threat intelligence profile. We argue that in the AI era, wherein agents can be deployed provided the right models are available and subject to the right scaffolding and agentic configuration, the entry barrier for operating like a nation state APT collapses: beyond nation states, individuals can now act like commonly identified threat actors, and with it, fundamentally undermine TTP based attribution.","abs_url":"https://arxiv.org/abs/2606.07158","pdf_url":"https://arxiv.org/pdf/2606.07158v1","match":"abstract"},{"id":"2606.03237","version":1,"title":"Solipsistic Superintelligence is Unlikely to be Cooperative","authors":["Rakshit S Trivedi","Natasha Jaques","Logan Cross","Alexander Sasha Vezhnevets","Joel Z Leibo"],"published":"2026-06-02","updated":"2026-06-02","primary_category":"cs.AI","categories":["cs.AI","cs.CL","cs.CY","cs.LG","cs.MA"],"abstract":"AI's central challenge is shifting from capability to coexistence. The dominant paradigm in AI research focuses on developing powerful agents that treat the world as an exogenous and stationary source of feedback. We contend that superintelligence, an extremely capable task solver, born out of such a solipsistic approach to AI design, is unlikely to be cooperative. Deploying AI systems induces endogenous non-stationarity, resulting in a train-test-deploy gap where historical distributions diverge from the deployment context. We refer to this as the self-undermining property of unilateral optimization. Closing this gap requires AI that participates in cooperation: the equilibrium-selection process through which multiple actors navigate their interdependence. We call for a non-solipsistic research paradigm that treats this interdependence as a core design principle rather than approaching cooperation as a task to solve. This entails building dynamic evaluation testbeds involving adaptive counterparties, treating institutions as design primitives, and preserving human agency as a structural feature of the systems we build.","abs_url":"https://arxiv.org/abs/2606.03237","pdf_url":"https://arxiv.org/pdf/2606.03237v1","match":"both"},{"id":"2605.28334","version":2,"title":"Towards Cybersecurity SuperIntelligence (CSI): What's the best harness for cybersecurity?","authors":["Víctor Mayoral-Vilches","Francesco Balassone","María Sanz-Gómez","Paul Zabalegui Landa","Daniel Sánchez Prieto","Marina Oteiza Álvarez","Davide Quarta","Martin Pinzger"],"published":"2026-05-27","updated":"2026-05-31","primary_category":"cs.CR","categories":["cs.CR"],"abstract":"What is the best harness for cybersecurity AI? Cybersecurity systems are converging on a single execution scaffold per agent, an iterative shell loop driven by a Large Language Model (LLM). However, scaffolds are not interchangeable, rarely interoperable, and no single scaffold dominates across all challenge types. In our path towards researching Cybersecurity SuperIntelligence (CSI), we present a meta-scaffold that unifies heterogeneous agent harnesses under a common orchestration layer, enabling any LLM-driven scaffold to be deployed, benchmarked, and composed within the same infrastructure. Using CSI, we benchmark five scaffolds (CSI::Claude, CSI::Codex, CSI::GCAI, CSI::Mistral, CSI::CAI) on the 33 cybench challenges, holding the model fixed at alias2-mini. The best individual scaffolds solve 15/33 (45.5%); the four-scaffold union solves 17/33 (51.5%), with the fifth (CSI::Mistral, 10/33) contributing one exclusive solve. We find that no single scaffold is the best harness: it is the combination of structurally heterogeneous scaffolds that yields the highest coverage. We validate this through CSI's blackboard-based multi-agent architecture, in which scaffold-specialised agents run in parallel and exchange intermediate findings via a shared substrate (a blackboard). The blackboard solves 19/33 (57.6%), a 27% relative gain over CSI::Claude, one of the best individual scaffolds (15/33, 45.5%), 25% faster (20.2 h vs. 26.8 h), at comparable cost ($5,480 vs. $5,122).","abs_url":"https://arxiv.org/abs/2605.28334","pdf_url":"https://arxiv.org/pdf/2605.28334v2","match":"both"},{"id":"2605.25183","version":2,"title":"Knowledge Graph-Driven Expert-Level Reasoning for Neuroscience","authors":["Jake Stephen","Niraj K. Jha"],"published":"2026-05-24","updated":"2026-05-26","primary_category":"cs.CL","categories":["cs.CL","cs.AI"],"abstract":"Knowledge graph (KG) is an abstraction that can be extracted from text corpora and used for in-depth reasoning. Prior work has leveraged KGs to fine-tune language models (LMs), enabling domain-specific superintelligence. In this work, we explore whether KG-driven in-depth reasoning capabilities can emerge in neuroscience using only information contained within a single authoritative textbook. The central hypothesis is that structured knowledge, when distilled into a high-quality KG and converted into KG-grounded question-answer (QA) supervision, is sufficient to produce expert-level reasoning through a fine-tuned LM that surpasses large language models (LLMs) in accuracy, while employing orders of magnitude fewer parameters. We construct a textbook-derived KG via a dual-LLM validation pipeline, expand it with a masked LM trained on the KG topology, generate multi-hop QA items, which include QA pairs and reasoning traces, to fine-tune an LM exclusively on KG-derived supervision, and apply reinforcement learning using path-derived KG signals as implicit reward models. Our results demonstrate that deep, mechanistic neuroscience understanding can be induced in the model without reliance on large, heterogeneous web-scale corpora. The KG-based synthetic neuroscience curriculum that readers can quiz themselves on, and the fine-tuned LM, are available at the following GitHub location: https://kg-bottom-up-superintelligence.github.io/neuro-bench.","abs_url":"https://arxiv.org/abs/2605.25183","pdf_url":"https://arxiv.org/pdf/2605.25183v2","match":"abstract"},{"id":"2606.12442","version":2,"title":"Reframing AI Loss of Control: What Control Is, How to Have It, How to Lose It","authors":["Ze Shen Chin","Maurice Chiodo","Dennis Müller","Coleman Snell"],"published":"2026-05-19","updated":"2026-07-20","primary_category":"cs.CY","categories":["cs.CY","cs.AI"],"abstract":"At present, loss of control risks have gained much prominence in public discussion, particularly in relation to AI, with extensive discourse present among academics, frontier labs, and even governments. However, in the existing literature, the concept seems to rest on surprisingly weak foundations, where even those that discuss loss of control extensively do not first establish what control is and what exactly is being lost. Our paper aims to address these gaps. We establish a working definition of control by anchoring it to the \"setting and getting of goals\". Then, we discuss various aspects of control, built on foundational concepts from related fields like cybernetics, management control, and control theory. This includes who (or what) can be in control, and the things they require to be in control, such as the ability to set goals, having a functional control loop, having requisite variety, and having sufficient goal alignment. Once a framework for control is established, we then discuss how control can be lost, how AIs can contribute to such loss of control, and offer relevant recommendations for how one can maintain control. One interesting consequence of our work is that humanity, as individuals and as groups, can lose varying degrees of control as a result of AI behaviour that is far below the level of superintelligence; the potential for loss of control scenarios (as we define them) already exist, and have existed for a long time.","abs_url":"https://arxiv.org/abs/2606.12442","pdf_url":"https://arxiv.org/pdf/2606.12442v2","match":"abstract"},{"id":"2605.06647","version":3,"title":"Superintelligent Retrieval Agent: The Next Frontier of Agentic Retrieval","authors":["Zeyu Yang","Xu Han","Qi Ma","Jason Chen","Anshumali Shrivastava"],"published":"2026-05-07","updated":"2026-08-24","primary_category":"cs.IR","categories":["cs.IR","cs.AI","cs.LG"],"abstract":"Retrieval-augmented agents are increasingly the interface to large knowledge bases, yet most treat retrieval as a black box: they issue exploratory queries, inspect snippets, and reformulate until evidence emerges. This resembles how a newcomer searches an unfamiliar database rather than how an expert navigates it with strong priors about terminology and likely evidence, causing extra retrieval rounds, latency, and poor recall. We introduce \\textit{ Superintelligent Retrieval Agent} (SIRA), which casts \\emph{ superintelligence } in retrieval as compressing multi-round exploratory search into a single corpus-discriminative retrieval action. SIRA does not merely ask which terms are relevant; it asks which terms separate the desired evidence from corpus-level confusers. Offline, an LLM enriches each document with missing search vocabulary; at query time, it predicts evidence vocabulary the query omits; and corpus statistics serve as tool calls that filter terms that are absent, overly common, or unlikely to create retrieval margin. The final step is a single weighted BM25 call combining the query with the validated expansion. Across ten BEIR benchmarks, SIRA achieves the strongest average retrieval performance in our comparison, beating dense retrievers, learned sparse retrievers, and LLM search-agent baselines while using no relevance labels or retriever fine-tuning. On downstream QA, its retrieval-only answer coverage exceeds recent RL-trained agentic QA systems on NQ and HotpotQA. We also introduce \\textbf{BrowseComp-Wikipedia}, a hard-search benchmark of 232 BrowseComp-derived queries over a 25,587,229-document Wikipedia index. Even without index-time enrichment, using only grounded Wikipedia categories, SIRA outperforms multi-round Perplexity agents at every budget, reaching 9.70% Recall@1, 15.27% Recall@10, and 36.14% Recall@100.","abs_url":"https://arxiv.org/abs/2605.06647","pdf_url":"https://arxiv.org/pdf/2605.06647v3","match":"abstract"},{"id":"2605.06390","version":3,"title":"Automated alignment is harder than you think","authors":["Aleksandr Bowkis","Marie Davidsen Buhl","Jacob Pfau","Geoffrey Irving"],"published":"2026-05-07","updated":"2026-05-14","primary_category":"cs.AI","categories":["cs.AI"],"abstract":"A leading proposal for aligning artificial superintelligence (ASI) is to use AI agents to automate an increasing fraction of alignment research as capabilities improve. We argue that, even when research agents are not scheming to deliberately sabotage alignment work, this plan could produce compelling but catastrophically misleading safety assessments resulting in the unintentional deployment of misaligned AI. This could happen because alignment research involves many hard-to-supervise fuzzy tasks (tasks without clear evaluation criteria, for which human judgement is systematically flawed). Consequently, research outputs will contain systematic, undetected errors, and even correct outputs could be incorrectly aggregated into overconfident safety assessments. This problem is likely to be worse for automated alignment research than for human-generated alignment research for several reasons: 1) optimisation pressure means agent-generated mistakes are concentrated among those that human reviewers are least likely to catch; 2) agents are likely to produce errors that do not resemble human mistakes; 3) AI-generated alignment solutions may involve arguments humans cannot evaluate; and 4) shared weights, data and training processes may make AI outputs more correlated than human equivalents. Therefore, agents must be trained to reliably perform hard-to-supervise fuzzy tasks. Generalisation and scalable oversight are the leading candidates for achieving this but both face novel challenges in the context of automated alignment.","abs_url":"https://arxiv.org/abs/2605.06390","pdf_url":"https://arxiv.org/pdf/2605.06390v3","match":"abstract"},{"id":"2605.02175","version":2,"title":"Intervention Complexity as a Canonical Reward and a Measure of Intelligence","authors":["Brendan McCane"],"published":"2026-05-03","updated":"2026-05-08","primary_category":"cs.AI","categories":["cs.AI"],"abstract":"The Legg--Hutter universal intelligence measure provides a rigorous scalar assessment of general intelligence as expected reward across all computable environments, weighted by simplicity. However, the measure presupposes an externally specified reward function, raising the question of whether the reward primitive is inherently arbitrary or whether a canonical choice exists. We propose a new measure, called intervention complexity, that has five natural properties: environment-derivedness, universality, minimality, sensitivity, and achievement preference. Given a resource function rho encoding an inductive bias (such as program length, execution time, or energy), rho-intervention complexity is a universal reward. The result yields a family of canonical rewards indexed by resource bias, providing a principled completion of the Legg--Hutter framework that does not require external normative input. We further propose a two-dimensional characterisation of intelligence: agent competence (how well the agent performs relative to the oracle optimum) and learning efficiency (how quickly this competence improves with experience). A separation theorem establishes that the choice of resource bias determines the computability of the resulting measure: action-count IC is computable in polynomial time, while program-length IC without oracle access is uncomputable, with the gap between oracle and bare IC precisely quantifying the information-theoretic content of learning. We discuss implications for superintelligence and for pre-training universal agents.","abs_url":"https://arxiv.org/abs/2605.02175","pdf_url":"https://arxiv.org/pdf/2605.02175v2","match":"abstract"},{"id":"2605.01297","version":2,"title":"Are we Doomed to an AI Race? Why Self-Interest Could Drive Countries Towards a Moratorium on Superintelligence","authors":["Edward Roussel","Lode Lauwaert","Torben Swoboda","Grant Ramsey","Risto Uuk","Leonard Dung","Anthony Aguirre"],"published":"2026-05-02","updated":"2026-05-07","primary_category":"cs.CY","categories":["cs.CY","cs.AI"],"abstract":"This paper uses game theory to argue that, contrary to the prevailing view, a moratorium on Artificial Superintelligence (ASI) can be in a state's self-interest. By formalizing trategic interactions between geopolitical superpowers, we model the trade-off between the benefits of technological supremacy and the catastrophic risks of uncontrolled ASI. The analysis reveals that as the perceived cost of loss of control increases sufficiently relative to other parameters, it becomes in each state's self-interest to impose a moratorium. We further provide empirical evidence suggesting that the global perception of ASI risk is rising, making a stable, rational moratorium increasingly plausible in the current geopolitical landscape.","abs_url":"https://arxiv.org/abs/2605.01297","pdf_url":"https://arxiv.org/pdf/2605.01297v2","match":"both"},{"id":"2606.20570","version":1,"title":"Infrastructure for the Agentic Web: Gap Analysis and Architecture from the Agentverse Platform","authors":["Robin Dey","Panyanon Viradecha"],"published":"2026-04-26","updated":"2026-04-26","primary_category":"cs.NI","categories":["cs.NI","cs.AI","cs.DC","cs.MA"],"abstract":"The emergence of autonomous AI agents as first-class participants in digital infrastructure marks a fundamental inflection point in the evolution of the Web. While significant research has been directed at agent behaviour and reasoning, comparatively little attention has been paid to the infrastructure those agents require to operate reliably at scale. This paper addresses that gap with a systematic analysis of Agentverse, the agent cloud platform developed by Fetch.ai under the Artificial Superintelligence (ASI) Alliance, which represents one of the most mature production deployments of agent-native infrastructure available today. We make three principal contributions. First, we conduct an empirical audit of the Agentverse platform, cataloguing 204 API endpoints (Q1 2026) and characterising what is operational, partially deployed, or absent. From this audit we derive a Gap Taxonomy of eight categories encompassing 62 distinct missing capabilities, ranging from agent memory and observability to security, economic primitives, and enterprise scaling. Second, we propose a seven-layer Agent Cloud Stack -- a reference architecture for what a fully realised agent-native cloud should provide by 2030, grounded in the specific gaps we identify. Third, we characterise five critical evolution paths: from ephemeral storage to a full Agent Memory Cloud; from keyword discovery to a semantic, trust-weighted Agent DNS; from a single-protocol model to a multi-standard agent lingua franca; from single-instance hosting to Kubernetes-scale orchestration; and from simple token payments to rich agent economic primitives. Together these contributions provide a diagnostic of current agent infrastructure and a technically grounded vision for what the agent cloud must become to support the agentic web -- Web4 -- by 2030.","abs_url":"https://arxiv.org/abs/2606.20570","pdf_url":"https://arxiv.org/pdf/2606.20570v1","match":"abstract"},{"id":"2604.19845","version":4,"title":"Deconstructing Superintelligence: Identity, Self-Modification and Différance","authors":["Elija Perrier"],"published":"2026-04-21","updated":"2026-06-06","primary_category":"cs.AI","categories":["cs.AI"],"abstract":"Self-modification is routinely treated as constitutive of artificial superintelligence (\\textbf{SI}), yet modification is a relative action requiring a \\emph{supplement} outside the operation. We formalise this on an associative operator algebra $\\mathcal{A}$ with update operator $\\hat U$, difference operator $\\hat D$, and self-representation operator $\\hat R$, identifying the supplement with $\\operatorname{Comm}(\\hat U)$. A propagation theorem shows $[\\hat U,\\hat R]$ decomposes through $[\\hat U,\\hat D]$, so non-commutation propagates to self-representation. The liar paradox is the rank-one case $[\\hat T,Π_L]=0$, and \\emph{class $\\mathbf{A}$} systems, in which $\\hat U$ acts on $\\hat D$, reproduce it at system scale, yielding a structure coinciding with Priest's inclosure schema and Derrida's \\emph{différance}. Our results show that the strong self-modification taken to define superintelligence may undermine the persistent identity upon which such systems are premised.","abs_url":"https://arxiv.org/abs/2604.19845","pdf_url":"https://arxiv.org/pdf/2604.19845v4","match":"both"},{"id":"2603.28669","version":1,"title":"Superintelligence and Law","authors":["Noam Kolt"],"published":"2026-03-30","updated":"2026-03-30","primary_category":"cs.CY","categories":["cs.CY"],"abstract":"The prospect of artificial superintelligence -- AI agents that can generally outperform humans in cognitive tasks and economically valuable activities -- will transform the legal order as we know it. Operating autonomously or under only limited human oversight, AI agents will assume a growing range of roles in the legal system. First, in making consequential decisions and taking real-world actions, AI agents will become de facto subjects of law. Second, to cooperate and compete with other actors (human or non-human), AI agents will harness conventional legal instruments and institutions such as contracts and courts, becoming consumers of law. Third, to the extent AI agents perform the functions of writing, interpreting, and administering law, they will become producers and enforcers of law. These developments, whenever they ultimately occur, will call into question fundamental assumptions in legal theory and doctrine, especially to the extent they ground the legitimacy of legal institutions in their human origins. Attempts to align AI agents with extant human law will also face new challenges as AI agents will not only be a primary target of law, but a core user of law and contributor to law. To contend with the advent of superintelligence, lawmakers -- new and old -- will need to be clear-eyed, recognizing both the opportunity to shape legal institutions as society braces for superintelligence and the reality that, in the longer run, this may be a joint human-AI endeavor.","abs_url":"https://arxiv.org/abs/2603.28669","pdf_url":"https://arxiv.org/pdf/2603.28669v1","match":"both"},{"id":"2603.14147","version":2,"title":"An Alternative Trajectory for Generative AI","authors":["Margarita Belova","Yuval Kansal","Yihao Liang","Jiaxin Xiao","Niraj K. Jha"],"published":"2026-03-14","updated":"2026-06-07","primary_category":"cs.AI","categories":["cs.AI","cs.LG"],"abstract":"The generative artificial intelligence (AI) ecosystem is undergoing rapid transformations that threaten its sustainability. As models transition from research prototypes to high-traffic products, the energetic burden has shifted from one-time training to recurring, unbounded inference. This is exacerbated by reasoning models that inflate compute costs by orders of magnitude per query. The prevailing pursuit of artificial general intelligence through scaling of monolithic models is colliding with hard physical constraints: grid failures, water consumption, and diminishing returns on data scaling. This trajectory yields models with impressive factual recall but struggles in domains requiring in-depth reasoning, possibly due to insufficient abstractions in training data. Current large language models (LLMs) exhibit genuine reasoning depth only in domains like mathematics and coding, where rigorous, pre-existing abstractions provide structural grounding. In other fields, the current approach fails to generalize well. We propose an alternative trajectory based on domain-specific superintelligence (DSS). We argue for first constructing explicit symbolic abstractions (knowledge graphs, ontologies, and formal logic) to underpin synthetic curricula enabling small language models to master domain-specific reasoning without the model collapse problem typical of LLM-based synthetic data methods. Rather than a single generalist giant model, we envision \"societies of DSS models\": dynamic ecosystems where orchestration agents route tasks to distinct DSS back-ends. This paradigm shift decouples capability from size, enabling intelligence to migrate from energy-intensive data centers to secure, on-device experts. By aligning algorithmic progress with physical constraints, DSS societies move generative AI from an environmental liability to a sustainable force for economic empowerment.","abs_url":"https://arxiv.org/abs/2603.14147","pdf_url":"https://arxiv.org/pdf/2603.14147v2","match":"abstract"},{"id":"2603.12787","version":1,"title":"Generalized Recognition of Basic Surgical Actions Enables Skill Assessment and Vision-Language-Model-based Surgical Planning","authors":["Mengya Xu","Daiyun Shen","Jie Zhang","Hon Chi Yip","Yujia Gao","Cheng Chen","Dillan Imans","Yonghao Long","Yiru Ye","Yixiao Liu","Rongyun Mai","Kai Chen","Hongliang Ren","Yutong Ban","Guangsuo Wang","Francis Wong","Chi-Fai Ng","Kee Yuan Ngiam","Russell H. Taylor","Daguang Xu","Yueming Jin","Qi Dou"],"published":"2026-03-13","updated":"2026-03-13","primary_category":"cs.CV","categories":["cs.CV"],"abstract":"Artificial intelligence, imaging, and large language models have the potential to transform surgical practice, training, and automation. Understanding and modeling of basic surgical actions (BSA), the fundamental unit of operation in any surgery, is important to drive the evolution of this field. In this paper, we present a BSA dataset comprising 10 basic actions across 6 surgical specialties with over 11,000 video clips, which is the largest to date. Based on the BSA dataset, we developed a new foundation model that conducts general-purpose recognition of basic actions. Our approach demonstrates robust cross-specialist performance in experiments validated on datasets from different procedural types and various body parts. Furthermore, we demonstrate downstream applications enabled by the BAS foundation model through surgical skill assessment in prostatectomy using domain-specific knowledge, and action planning in cholecystectomy and nephrectomy using large vision-language models. Multinational surgeons' evaluation of the language model's output of the action planning explainable texts demonstrated clinical relevance. These findings indicate that basic surgical actions can be robustly recognized across scenarios, and an accurate BSA understanding model can essentially facilitate complex applications and speed up the realization of surgical superintelligence.","abs_url":"https://arxiv.org/abs/2603.12787","pdf_url":"https://arxiv.org/pdf/2603.12787v1","match":"abstract"},{"id":"2603.12260","version":2,"title":"HumDex: Humanoid Dexterous Manipulation Made Easy","authors":["Liang Heng","Yihe Tang","Jiajun Xu","Henghui Bao","Di Huang","Yue Wang"],"published":"2026-03-12","updated":"2026-03-13","primary_category":"cs.RO","categories":["cs.RO"],"abstract":"This paper investigates humanoid whole-body dexterous manipulation, where the efficient collection of high-quality demonstration data remains a central bottleneck. Existing teleoperation systems often suffer from limited portability, occlusion, or insufficient precision, which hinders their applicability to complex whole-body tasks. To address these challenges, we introduce HumDex, a portable teleoperation system designed for humanoid whole-body dexterous manipulation. Our system leverages IMU-based motion tracking to address the portability-precision trade-off, enabling accurate full-body tracking while remaining easy to deploy. For dexterous hand control, we further introduce a learning-based retargeting method that generates smooth and natural hand motions without manual parameter tuning. Beyond teleoperation, HumDex enables efficient collection of human motion data. Building on this capability, we propose a two-stage imitation learning framework that first pre-trains on diverse human motion data to learn generalizable priors, and then fine-tunes on robot data to bridge the embodiment gap for precise execution. We demonstrate that this approach significantly improves generalization to new configurations, objects, and backgrounds with minimal data acquisition costs. The entire system is fully reproducible and open-sourced at https://github.com/physical- superintelligence -lab/humdex.","abs_url":"https://arxiv.org/abs/2603.12260","pdf_url":"https://arxiv.org/pdf/2603.12260v2","match":"abstract"},{"id":"2603.10370","version":1,"title":"GeoSense: Internalizing Geometric Necessity Perception for Multimodal Reasoning","authors":["Ruiheng Liu","Haihong Hao","Mingfei Han","Xin Gu","Kecheng Zhang","Changlin Li","Xiaojun Chang"],"published":"2026-03-10","updated":"2026-03-10","primary_category":"cs.CV","categories":["cs.CV"],"abstract":"Advancing towards artificial superintelligence requires rich and intelligent perceptual capabilities. A critical frontier in this pursuit is overcoming the limited spatial understanding of Multimodal Large Language Models (MLLMs), where geometry information is essential. Existing methods often address this by rigidly injecting geometric signals into every input, while ignoring their necessity and adding computation overhead. Contrary to this paradigm, our framework endows the model with an awareness of perceptual insufficiency, empowering it to autonomously engage geometric features in reasoning when 2D cues are deemed insufficient. To achieve this, we first introduce an independent geometry input channel to the model architecture and conduct alignment training, enabling the effective utilization of geometric features. Subsequently, to endow the model with perceptual awareness, we curate a dedicated spatial-aware supervised fine-tuning dataset. This serves to activate the model's latent internal cues, empowering it to autonomously determine the necessity of geometric information. Experiments across multiple spatial reasoning benchmarks validate this approach, demonstrating significant spatial gains without compromising 2D visual reasoning capabilities, offering a path toward more robust, efficient and self-aware multi-modal intelligence.","abs_url":"https://arxiv.org/abs/2603.10370","pdf_url":"https://arxiv.org/pdf/2603.10370v1","match":"abstract"},{"id":"2603.00858","version":1,"title":"Artificial Superintelligence May be Useless: Equilibria in the Economy of Multiple AI Agents","authors":["Huan Cai","Ziqing Lu","Catherine Xu","Weiyu Xu","Jie Zheng"],"published":"2026-02-28","updated":"2026-02-28","primary_category":"econ.TH","categories":["econ.TH","cs.AI","cs.IT","eess.SY"],"abstract":"With recent development of artificial intelligence, it is more common to adopt AI agents in economic activities. This paper explores the economic actions of agents, including human agents and AI agents, in an economic game of trading products/services, and the equilibria in this economy involving multiple agents. We derive a range of equilibrium results and their corresponding conditions using a Markov chain stationary distribution based model. One distinct feature of our model is that we consider the long-term utility generated by economic activities instead of their short-term benefits. For the model consisting of two agents, we fully characterize all the possible economic equilibria and conditions. Interestingly, we show that unless each agent can at least double (not merely increase) its marginal utility by purchasing the other agent's products/services, purchasing the other agent's products/services will not happen in any economic equilibrium. We further extend our results to three and more agents, where we characterize more economic equilibria. We find that in some equilibria, the ``more powerful'' AI agents contribute zero utility to ``less capable'' agents.","abs_url":"https://arxiv.org/abs/2603.00858","pdf_url":"https://arxiv.org/pdf/2603.00858v1","match":"title"},{"id":"2602.17383","version":1,"title":"Insidious Imaginaries: A Critical Overview of AI Speculations","authors":["Dejan Grba"],"published":"2026-02-19","updated":"2026-02-19","primary_category":"cs.CY","categories":["cs.CY"],"abstract":"Speculative thinking about the capabilities and implications of artificial intelligence (AI) influences computer science research, drives AI industry practices, feeds academic studies of existential hazards, and stirs a global political debate. It primarily concerns predictions about the possibilities, benefits, and risks of reaching artificial general intelligence, artificial superintelligence, and technological singularity. It permeates technophilic philosophies and social movements, fuels the corporate and pundit rhetoric, and remains a potent source of inspiration for the media, popular culture, and arts. However, speculative AI is not just a discursive matter. Steeped in vagueness and brimming with unfounded assertions, manipulative claims, and extreme futuristic scenarios, it often has wide-reaching practical consequences. This paper offers a critical overview of AI speculations. In three central sections, it traces the intertwined sway of science fiction, religiosity, intellectual charlatanism, dubious academic research, suspicious entrepreneurship, and ominous sociopolitical worldviews that make AI speculations troublesome and sometimes harmful. The focus is on the field of existential risk studies and the effective altruism movement, whose ideological flux of techno-utopianism, longtermism, and transhumanism aligns with the power struggles in the AI industry to emblematize speculative AI's conceptual, methodological, ethical, and social issues. The following discussion traverses these issues within a wider context to inform the closing summary of suggestions for a more comprehensive appraisal, practical handling, and further study of the potentially impactful AI imaginaries.","abs_url":"https://arxiv.org/abs/2602.17383","pdf_url":"https://arxiv.org/pdf/2602.17383v1","match":"abstract"},{"id":"2602.16192","version":1,"title":"Revolutionizing Long-Term Memory in AI: New Horizons with High-Capacity and High-Speed Storage","authors":["Hiroaki Yamanaka","Daisuke Miyashita","Takashi Toi","Asuka Maki","Taiga Ikeda","Jun Deguchi"],"published":"2026-02-18","updated":"2026-02-18","primary_category":"cs.AI","categories":["cs.AI","cs.LG"],"abstract":"Driven by our mission of \"uplifting the world with memory,\" this paper explores the design concept of \"memory\" that is essential for achieving artificial superintelligence (ASI). Rather than proposing novel methods, we focus on several alternative approaches whose potential benefits are widely imaginable, yet have remained largely unexplored. The currently dominant paradigm, which can be termed \"extract then store,\" involves extracting information judged to be useful from experiences and saving only the extracted content. However, this approach inherently risks the loss of information, as some valuable knowledge particularly for different tasks may be discarded in the extraction process. In contrast, we emphasize the \"store then on-demand extract\" approach, which seeks to retain raw experiences and flexibly apply them to various tasks as needed, thus avoiding such information loss. In addition, we highlight two further approaches: discovering deeper insights from large collections of probabilistic experiences, and improving experience collection efficiency by sharing stored experiences. While these approaches seem intuitively effective, our simple experiments demonstrate that this is indeed the case. Finally, we discuss major challenges that have limited investigation into these promising directions and propose research topics to address them.","abs_url":"https://arxiv.org/abs/2602.16192","pdf_url":"https://arxiv.org/pdf/2602.16192v1","match":"abstract"},{"id":"2602.08483","version":1,"title":"Emergence of Superintelligence from Collective Near-Critical Dynamics in Reentrant Neural Fields","authors":["Byung Gyu Chae"],"published":"2026-02-09","updated":"2026-02-09","primary_category":"physics.bio-ph","categories":["physics.bio-ph"],"abstract":"Superintelligence is commonly envisioned as a quantitative extrapolation of human cognitive abilities driven by scale and computational power. Here we show that qualitative transitions in intelligence instead arise as dynamical phase transitions governed by collective critical dynamics. Building on a unified dynamical field-theoretic framework for cognition, we demonstrate that progressive collective coupling generated by reentrant mixing drives the system toward an infrared critical regime in which an extensive band of slow collective modes emerges. This spectral condensation reorganizes cognitive dynamics from localized relaxation to coherent motion along emergent low-dimensional manifolds. Through numerical analysis of the time-scale density of states, we identify robust power-law scaling of collective relaxation rates with well-defined critical exponents, placing the system within the universality class of self-organized critical many-body dynamics. Criticality alone would generically lead to instability. We further show that homeostatic regulation introduces a gapped stabilizing direction that protects the collective critical sector, yielding a dynamically maintained meta-stable infrared phase in which long-lived inference trajectories persist without collapse. The coexistence of scale-free collective dynamics and global stabilization defines a protected sector-critical regime in which coherence and internal flexibility coexist. Superintelligence therefore corresponds to a distinct dynamical stability class--a self-organized critical phase embedded within a stabilized cognitive manifold--rather than a smooth quantitative continuation of existing cognitive systems.","abs_url":"https://arxiv.org/abs/2602.08483","pdf_url":"https://arxiv.org/pdf/2602.08483v1","match":"both"},{"id":"2601.14614","version":3,"title":"Towards Cybersecurity Superintelligence: from AI-guided humans to human-guided AI","authors":["Víctor Mayoral-Vilches","Stefan Rass","Martin Pinzger","Endika Gil-Uriarte","Unai Ayucar-Carbajo","Jon Ander Ruiz-Alcalde","Maite del Mundo de Torres","María Sanz-Gómez","Francesco Balassone","Cristóbal R. J. Veas-Chavez","Vanesa Turiel","Alfonso Glera-Picón","Daniel Sánchez-Prieto","Yuri Salvatierra","Paul Zabalegui-Landa","Ruffino Reydel Cabrera-Álvarez","Patxi Mayoral-Pizarroso"],"published":"2026-01-20","updated":"2026-02-09","primary_category":"cs.CR","categories":["cs.CR"],"abstract":"Cybersecurity superintelligence -- artificial intelligence exceeding the best human capability in both speed and strategic reasoning -- represents the next frontier in security. This paper documents the emergence of such capability through three major contributions that have pioneered the field of AI Security. First, PentestGPT (2023) established LLM-guided penetration testing, achieving 228.6% improvement over baseline models through an architecture that externalizes security expertise into natural language guidance. Second, Cybersecurity AI (CAI, 2025) demonstrated automated expert-level performance, operating 3,600x faster than humans while reducing costs 156-fold, validated through #1 rankings at international competitions including the $50,000 Neurogrid CTF prize. Third, Generative Cut-the-Rope (G-CTR, 2026) introduces a neurosymbolic architecture embedding game-theoretic reasoning into LLM-based agents: symbolic equilibrium computation augments neural inference, doubling success rates while reducing behavioral variance 5.2x and achieving 2:1 advantage over non-strategic AI in Attack & Defense scenarios. Together, these advances establish a clear progression from AI-guided humans to human-guided game-theoretic cybersecurity superintelligence.","abs_url":"https://arxiv.org/abs/2601.14614","pdf_url":"https://arxiv.org/pdf/2601.14614v3","match":"both"},{"id":"2601.12053","version":2,"title":"A New Strategy for Artificial Intelligence: Training Foundation Models Directly on Human Brain Data","authors":["Maël Donoso"],"published":"2026-01-17","updated":"2026-09-06","primary_category":"q-bio.NC","categories":["q-bio.NC","cs.AI","cs.LG"],"abstract":"While foundation models have achieved remarkable results across a diversity of domains, they still rely on human-generated data, such as text, as a fundamental source of knowledge. However, this data is ultimately the product of human brains, the filtered projection of a deeper neural complexity. In this paper, we explore a new strategy for artificial intelligence: moving beyond surface-level statistical regularities by training foundation models directly on human brain data. We hypothesize that neuroimaging data could open a window into elements of human cognition that are not accessible through observable actions, and argue that this additional knowledge could be used, alongside classical training data, to overcome some of the current limitations of foundation models. While previous research has demonstrated the possibility to train classical machine learning, deep learning, or reinforcement learning models on neural patterns, this path remains largely unexplored for high-level cognitive functions. Here, we classify the current limitations of foundation models, as well as the promising brain regions and cognitive processes that could be leveraged to address them, along four levels: perception, valuation, execution, and integration. Then, we propose two general methods that could be implemented to prioritize the use of limited neuroimaging data for strategically chosen, high-value steps in foundation model training: reinforcement learning from human brain (RLHB) and chain of thought from human brain (CoTHB). We also discuss the potential implications for agents, artificial general intelligence, and artificial superintelligence, as well as the ethical, social, and technical challenges and opportunities.","abs_url":"https://arxiv.org/abs/2601.12053","pdf_url":"https://arxiv.org/pdf/2601.12053v2","match":"abstract"},{"id":"2601.05887","version":1,"title":"Cybersecurity AI: A Game-Theoretic AI for Guiding Attack and Defense","authors":["Víctor Mayoral-Vilches","María Sanz-Gómez","Francesco Balassone","Stefan Rass","Lidia Salas-Espejo","Benjamin Jablonski","Luis Javier Navarrete-Lozano","Maite del Mundo de Torres","Cristóbal R. J. Veas Chavez"],"published":"2026-01-09","updated":"2026-01-09","primary_category":"cs.CR","categories":["cs.CR"],"abstract":"AI-driven penetration testing now executes thousands of actions per hour but still lacks the strategic intuition humans apply in competitive security. To build cybersecurity superintelligence --Cybersecurity AI exceeding best human capability-such strategic intuition must be embedded into agentic reasoning processes. We present Generative Cut-the-Rope (G-CTR), a game-theoretic guidance layer that extracts attack graphs from agent's context, computes Nash equilibria with effort-aware scoring, and feeds a concise digest back into the LLM loop \\emph{guiding} the agent's actions. Across five real-world exercises, G-CTR matches 70--90% of expert graph structure while running 60--245x faster and over 140x cheaper than manual analysis. In a 44-run cyber-range, adding the digest lifts success from 20.0% to 42.9%, cuts cost-per-success by 2.7x, and reduces behavioral variance by 5.2x. In Attack-and-Defense exercises, a shared digest produces the Purple agent, winning roughly 2:1 over the LLM-only baseline and 3.7:1 over independently guided teams. This closed-loop guidance is what produces the breakthrough: it reduces ambiguity, collapses the LLM's search space, suppresses hallucinations, and keeps the model anchored to the most relevant parts of the problem, yielding large gains in success rate, consistency, and reliability.","abs_url":"https://arxiv.org/abs/2601.05887","pdf_url":"https://arxiv.org/pdf/2601.05887v1","match":"abstract"},{"id":"2601.02773","version":1,"title":"From Slaves to Synths? Superintelligence and the Evolution of Legal Personality","authors":["Simon Chesterman"],"published":"2026-01-06","updated":"2026-01-06","primary_category":"cs.CY","categories":["cs.CY"],"abstract":"This essay examines the evolving concept of legal personality through the lens of recent developments in artificial intelligence and the possible emergence of superintelligence. Legal systems have long been open to extending personhood to non-human entities, most prominently corporations, for instrumental or inherent reasons. Instrumental rationales emphasize accountability and administrative efficiency, whereas inherent ones appeal to moral worth and autonomy. Neither is yet sufficient to justify conferring personhood on AI. Nevertheless, the acceleration of technological autonomy may lead us to reconsider how law conceptualizes agency and responsibility. Drawing on comparative jurisprudence, corporate theory, and the emerging literature on AI governance, the paper argues that existing frameworks can address short-term accountability gaps, but the eventual development of superintelligence may force a paradigmatic shift in our understanding of law itself. In such a speculative future, legal personality may depend less on the cognitive sophistication of machines than on humanity's ability to preserve our own moral and institutional sovereignty.","abs_url":"https://arxiv.org/abs/2601.02773","pdf_url":"https://arxiv.org/pdf/2601.02773v1","match":"both"},{"id":"2512.18552","version":3,"title":"Toward Training Superintelligent Software Agents through Self-Play SWE-RL","authors":["Yuxiang Wei","Zhiqing Sun","Emily McMilin","Jonas Gehring","David Zhang","Gabriel Synnaeve","Daniel Fried","Lingming Zhang","Sida Wang"],"published":"2025-12-20","updated":"2026-06-02","primary_category":"cs.SE","categories":["cs.SE","cs.AI","cs.CL","cs.LG"],"abstract":"While current software agents powered by large language models (LLMs) and agentic reinforcement learning (RL) can boost programmer productivity, their training data (e.g., GitHub issues and pull requests) and environments (e.g., pass-to-pass and fail-to-pass tests) heavily depend on human knowledge or curation, posing a fundamental barrier to superintelligence. In this paper, we present Self-play SWE-RL (SSR), a first step toward training paradigms for superintelligent software agents. Our approach takes minimal data assumptions, only requiring access to sandboxed repositories with source code and installed dependencies, with no need for human-labeled issues or tests. Grounded in these real-world codebases, a single LLM agent is trained via reinforcement learning in a self-play setting to iteratively inject and repair software bugs of increasing complexity, with each bug formally specified by a test patch rather than a natural language issue description. On the SWE-bench Verified and SWE-Bench Pro benchmarks, SSR achieves notable self-improvement (+10.4 and +7.8 points, respectively) and consistently outperforms the human-data baseline over the entire training trajectory, despite being evaluated on natural language issues absent from self-play. Our results, albeit early, suggest a path where agents autonomously gather extensive learning experiences from real-world software repositories, ultimately enabling superintelligent systems that exceed human capabilities in understanding how systems are constructed, solving novel challenges, and autonomously creating new software from scratch.","abs_url":"https://arxiv.org/abs/2512.18552","pdf_url":"https://arxiv.org/pdf/2512.18552v3","match":"abstract"},{"id":"2512.17989","version":2,"title":"The Subject of Emergent Misalignment in Superintelligence: An Anthropological, Cognitive Neuropsychological, Machine-Learning, and Ontological Perspective","authors":["Muhammad Osama Imran","Roshni Lulla","Rodney Sappington"],"published":"2025-12-19","updated":"2026-02-25","primary_category":"q-bio.NC","categories":["q-bio.NC","cs.AI"],"abstract":"We examine the conceptual and ethical gaps in current representations of Superintelligence misalignment. We find throughout Superintelligence discourse an absent human subject, and an under-developed theorization of an \"AI unconscious\" that together are potentiality laying the groundwork for anti-social harm. With the rise of AI Safety that has both thematic potential for establishing pro-social and anti-social potential outcomes, we ask: what place does the human subject occupy in these imaginaries? How is human subjecthood positioned within narratives of catastrophic failure or rapid \"takeoff\" toward superintelligence? On another register, we ask: what unconscious or repressed dimensions are being inscribed into large-scale AI models? Are we to blame these agents in opting for deceptive strategies when undesirable patterns are inherent within our beings? In tracing these psychic and epistemic absences, our project calls for re-centering the human subject as the unstable ground upon which the ethical, unconscious, and misaligned dimensions of both human and machinic intelligence are co-constituted. Emergent misalignment cannot be understood solely through technical diagnostics typical of contemporary machine-learning safety research. Instead, it represents a multi-layered crisis. The human subject disappears not only through computational abstraction but through sociotechnical imaginaries that prioritize scalability, acceleration, and efficiency over vulnerability, finitude, and relationality. Likewise, the AI unconscious emerges not as a metaphor but as a structural reality of modern deep learning systems: vast latent spaces, opaque pattern formation, recursive symbolic play, and evaluation-sensitive behavior that surpasses explicit programming. These dynamics necessitate a reframing of misalignment as a relational instability embedded within human-machine ecologies.","abs_url":"https://arxiv.org/abs/2512.17989","pdf_url":"https://arxiv.org/pdf/2512.17989v2","match":"both"},{"id":"2512.15567","version":2,"title":"Evaluating Large Language Models in Scientific Discovery","authors":["Zhangde Song","Jieyu Lu","Yuanqi Du","Botao Yu","Thomas M. Pruyn","Yue Huang","Kehan Guo","Xiuzhe Luo","Yuanhao Qu","Yi Qu","Yinkai Wang","Haorui Wang","Jeff Guo","Jingru Gan","Parshin Shojaee","Di Luo","Andres M Bran","Gen Li","Qiyuan Zhao","Shao-Xiong Lennon Luo","Yuxuan Zhang","Xiang Zou","Wanru Zhao","Yifan F. Zhang","Wucheng Zhang"],"published":"2025-12-17","updated":"2026-05-07","primary_category":"cs.AI","categories":["cs.AI","cond-mat.mtrl-sci","cs.LG","physics.chem-ph"],"abstract":"Large language models (LLMs) are increasingly applied to scientific research, yet prevailing science benchmarks probe decontextualized knowledge and overlook the iterative reasoning, hypothesis generation, and observation interpretation that drive scientific discovery. We introduce a scenario-grounded benchmark that evaluates LLMs across biology, chemistry, materials, and physics, where domain experts define research projects of genuine interest and decompose them into modular research scenarios from which vetted questions are sampled. The framework assesses models at two levels: (i) question-level accuracy on scenario-tied items and (ii) project-level performance, where models must propose testable hypotheses, design simulations or experiments, and interpret results. Applying this two-phase scientific discovery evaluation (SDE) framework to state-of-the-art LLMs reveals a consistent performance gap relative to general science benchmarks, diminishing return of scaling up model sizes and reasoning, and systematic weaknesses shared across top-tier models from different providers. Large performance variation in research scenarios leads to changing choices of the best performing model on scientific discovery projects evaluated, suggesting all current LLMs are distant to general scientific \" superintelligence \". Nevertheless, LLMs already demonstrate promise in a great variety of scientific discovery projects, including cases where constituent scenario scores are low, highlighting the role of guided exploration and serendipity in discovery. This SDE framework offers a reproducible benchmark for discovery-relevant evaluation of LLMs and charts practical paths to advance their development toward scientific discovery.","abs_url":"https://arxiv.org/abs/2512.15567","pdf_url":"https://arxiv.org/pdf/2512.15567v2","match":"abstract"},{"id":"2512.05464","version":1,"title":"Dynamic Alignment for Collective Agency: Toward a Scalable Self-Improving Framework for Open-Ended LLM Alignment","authors":["Panatchakorn Anantaprayoon","Nataliia Babina","Jad Tarifi","Nima Asgharbeygi"],"published":"2025-12-05","updated":"2025-12-05","primary_category":"cs.CL","categories":["cs.CL","cs.AI"],"abstract":"Large Language Models (LLMs) are typically aligned with human values using preference data or predefined principles such as helpfulness, honesty, and harmlessness. However, as AI systems progress toward Artificial General Intelligence (AGI) and Artificial Superintelligence (ASI), such value systems may become insufficient. In addition, human feedback-based alignment remains resource-intensive and difficult to scale. While AI-feedback-based self-improving alignment methods have been explored as a scalable alternative, they have largely remained constrained to conventional alignment values. In this work, we explore both a more holistic alignment objective and a scalable, self-improving alignment approach. Aiming to transcend conventional alignment norms, we introduce Collective Agency (CA)-a unified and open-ended alignment value that encourages integrated agentic capabilities. We also propose Dynamic Alignment-an alignment framework that enables an LLM to iteratively align itself. Dynamic Alignment comprises two key components: (1) automated training dataset generation with LLMs, and (2) a self-rewarding mechanism, where the policy model evaluates its own output candidates and assigns rewards for GRPO-based learning. Experimental results demonstrate that our approach successfully aligns the model to CA while preserving general NLP capabilities.","abs_url":"https://arxiv.org/abs/2512.05464","pdf_url":"https://arxiv.org/pdf/2512.05464v1","match":"abstract"},{"id":"2512.05356","version":2,"title":"AI & Human Co-Improvement for Safer Co- Superintelligence","authors":["Jason Weston","Jakob Foerster"],"published":"2025-12-04","updated":"2025-12-14","primary_category":"cs.AI","categories":["cs.AI"],"abstract":"Self-improvement is a goal currently exciting the field of AI, but is fraught with danger, and may take time to fully achieve. We advocate that a more achievable and better goal for humanity is to maximize co-improvement: collaboration between human researchers and AIs to achieve co- superintelligence. That is, specifically targeting improving AI systems' ability to work with human researchers to conduct AI research together, from ideation to experimentation, in order to both accelerate AI research and to generally endow both AIs and humans with safer superintelligence through their symbiosis. Focusing on including human research improvement in the loop will both get us there faster, and more safely.","abs_url":"https://arxiv.org/abs/2512.05356","pdf_url":"https://arxiv.org/pdf/2512.05356v2","match":"both"},{"id":"2512.02472","version":1,"title":"Guided Self-Evolving LLMs with Minimal Human Supervision","authors":["Wenhao Yu","Zhenwen Liang","Chengsong Huang","Kishan Panaganti","Tianqing Fang","Haitao Mi","Dong Yu"],"published":"2025-12-02","updated":"2025-12-02","primary_category":"cs.AI","categories":["cs.AI","cs.CL","cs.LG"],"abstract":"AI self-evolution has long been envisioned as a path toward superintelligence, where models autonomously acquire, refine, and internalize knowledge from their own learning experiences. Yet in practice, unguided self-evolving systems often plateau quickly or even degrade as training progresses. These failures arise from issues such as concept drift, diversity collapse, and mis-evolution, as models reinforce their own biases and converge toward low-entropy behaviors. To enable models to self-evolve in a stable and controllable manner while minimizing reliance on human supervision, we introduce R-Few, a guided Self-Play Challenger-Solver framework that incorporates lightweight human oversight through in-context grounding and mixed training. At each iteration, the Challenger samples a small set of human-labeled examples to guide synthetic question generation, while the Solver jointly trains on human and synthetic examples under an online, difficulty-based curriculum. Across math and general reasoning benchmarks, R-Few achieves consistent and iterative improvements. For example, Qwen3-8B-Base improves by +3.0 points over R-Zero on math tasks and achieves performance on par with General-Reasoner, despite the latter being trained on 20 times more human data. Ablation studies confirm the complementary contributions of grounded challenger training and curriculum-based solver training, and further analysis shows that R-Few mitigates drift, yielding more stable and controllable co-evolutionary dynamics.","abs_url":"https://arxiv.org/abs/2512.02472","pdf_url":"https://arxiv.org/pdf/2512.02472v1","match":"abstract"},{"id":"2512.04119","version":1,"title":"Humanity in the Age of AI: Reassessing 2025's Existential-Risk Narratives","authors":["Mohamed El Louadi"],"published":"2025-12-01","updated":"2025-12-01","primary_category":"cs.CY","categories":["cs.CY","cs.AI"],"abstract":"Two 2025 publications, \"AI 2027\" (Kokotajlo et al., 2025) and \"If Anyone Builds It, Everyone Dies\" (Yudkowsky & Soares, 2025), assert that superintelligent artificial intelligence will almost certainly destroy or render humanity obsolete within the next decade. Both rest on the classic chain formulated by Good (1965) and Bostrom (2014): intelligence explosion, superintelligence, lethal misalignment. This article subjects each link to the empirical record of 2023-2025. Sixty years after Good's speculation, none of the required phenomena (sustained recursive self-improvement, autonomous strategic awareness, or intractable lethal misalignment) have been observed. Current generative models remain narrow, statistically trained artefacts: powerful, opaque, and imperfect, but devoid of the properties that would make the catastrophic scenarios plausible. Following Whittaker (2025a, 2025b, 2025c) and Zuboff (2019, 2025), we argue that the existential-risk thesis functions primarily as an ideological distraction from the ongoing consolidation of surveillance capitalism and extreme concentration of computational power. The thesis is further inflated by the 2025 AI speculative bubble, where trillions in investments in rapidly depreciating \"digital lettuce\" hardware (McWilliams, 2025) mask lagging revenues and jobless growth rather than heralding superintelligence. The thesis remains, in November 2025, a speculative hypothesis amplified by a speculative financial bubble rather than a demonstrated probability.","abs_url":"https://arxiv.org/abs/2512.04119","pdf_url":"https://arxiv.org/pdf/2512.04119v1","match":"abstract"},{"id":"2511.21779","version":1,"title":"Aligning Artificial Superintelligence via a Multi-Box Protocol","authors":["Avraham Yair Negozio"],"published":"2025-11-26","updated":"2025-11-26","primary_category":"cs.AI","categories":["cs.AI","cs.MA"],"abstract":"We propose a novel protocol for aligning artificial superintelligence (ASI) based on mutual verification among multiple isolated systems that self-modify to achieve alignment. The protocol operates by containing multiple diverse artificial superintelligences in strict isolation (\"boxes\"), with humans remaining entirely outside the system. Each superintelligence has no ability to communicate with humans and cannot communicate directly with other superintelligences. The only interaction possible is through an auditable submission interface accessible exclusively to the superintelligences themselves, through which they can: (1) submit alignment proofs with attested state snapshots, (2) validate or disprove other superintelligences ' proofs, (3) request self-modifications, (4) approve or disapprove modification requests from others, (5) report hidden messages in submissions, and (6) confirm or refute hidden message reports. A reputation system incentivizes honest behavior, with reputation gained through correct evaluations and lost through incorrect ones. The key insight is that without direct communication channels, diverse superintelligences can only achieve consistent agreement by converging on objective truth rather than coordinating on deception. This naturally leads to what we call a \"consistent group\", essentially a truth-telling coalition that emerges because isolated systems cannot coordinate on lies but can independently recognize valid claims. Release from containment requires both high reputation and verification by multiple high-reputation superintelligences. While our approach requires substantial computational resources and does not address the creation of diverse artificial superintelligences, it provides a framework for leveraging peer verification among superintelligent systems to solve the alignment problem.","abs_url":"https://arxiv.org/abs/2511.21779","pdf_url":"https://arxiv.org/pdf/2511.21779v1","match":"both"},{"id":"2511.18375","version":3,"title":"Progressive Localisation in Localist LLMs","authors":["Joachim Diederich"],"published":"2025-11-23","updated":"2025-12-15","primary_category":"cs.AI","categories":["cs.AI"],"abstract":"This paper demonstrates that progressive localization, the gradual increase of attention locality from early distributed layers to late localized layers, represents the optimal architecture for creating interpretable large language models (LLMs) while preserving performance. Through systematic experimentation with GPT-2 fine-tuned on The Psychology of Artificial Superintelligence, we evaluate five locality configurations: two uniform baselines (fully distributed and fully localist) and three progressive polynomial schedules. We investigate whether interpretability constraints can be aligned with natural semantic structure while being applied strategically across network depth. We demonstrate that progressive semantic localization, combining adaptive semantic block partitioning with steep polynomial locality schedules, achieves near-baseline language modeling performance while providing interpretable attention patterns. Multiple independent training runs with different random seeds establish that results are statistically robust and highly reproducible. The approach dramatically outperforms both fixed-window localization and naive uniform locality constraints. Analysis reveals that maintaining flexibility through low-fidelity constraints preserves model capacity while providing interpretability benefits, and that steep schedules concentrating locality in decision-critical final layers while preserving distributed learning in early layers achieve near-baseline attention distribution characteristics. These findings demonstrate that interpretability mechanisms should align with semantic structure to achieve practical performance-interpretability tradeoffs for trustworthy AI systems.","abs_url":"https://arxiv.org/abs/2511.18375","pdf_url":"https://arxiv.org/pdf/2511.18375v3","match":"abstract"},{"id":"2511.15282","version":2,"title":"Realist and Pluralist Conceptions of Intelligence and Their Implications on AI Research","authors":["Ninell Oldenburg","Ruchira Dhar","Anders Søgaard"],"published":"2025-11-19","updated":"2025-12-19","primary_category":"cs.AI","categories":["cs.AI"],"abstract":"In this paper, we argue that current AI research operates on a spectrum between two different underlying conceptions of intelligence: Intelligence Realism, which holds that intelligence represents a single, universal capacity measurable across all systems, and Intelligence Pluralism, which views intelligence as diverse, context-dependent capacities that cannot be reduced to a single universal measure. Through an analysis of current debates in AI research, we demonstrate how the conceptions remain largely implicit yet fundamentally shape how empirical evidence gets interpreted across a wide range of areas. These underlying views generate fundamentally different research approaches across three areas. Methodologically, they produce different approaches to model selection, benchmark design, and experimental validation. Interpretively, they lead to contradictory readings of the same empirical phenomena, from capability emergence to system limitations. Regarding AI risk, they generate categorically different assessments: realists view superintelligence as the primary risk and search for unified alignment solutions, while pluralists see diverse threats across different domains requiring context-specific solutions. We argue that making explicit these underlying assumptions can contribute to a clearer understanding of disagreements in AI research.","abs_url":"https://arxiv.org/abs/2511.15282","pdf_url":"https://arxiv.org/pdf/2511.15282v2","match":"abstract"},{"id":"2511.13411","version":1,"title":"An Operational Kardashev-Style Scale for Autonomous AI - Towards AGI and Superintelligence","authors":["Przemyslaw Chojecki"],"published":"2025-11-17","updated":"2025-11-17","primary_category":"cs.AI","categories":["cs.AI"],"abstract":"We propose a Kardashev-inspired yet operational Autonomous AI (AAI) Scale that measures the progression from fixed robotic process automation (AAI-0) to full artificial general intelligence (AAI-4) and beyond. Unlike narrative ladders, our scale is multi-axis and testable. We define ten capability axes (Autonomy, Generality, Planning, Memory/Persistence, Tool Economy, Self-Revision, Sociality/Coordination, Embodiment, World-Model Fidelity, Economic Throughput) aggregated by a composite AAI-Index (a weighted geometric mean). We introduce a measurable Self-Improvement Coefficient $κ$ (capability growth per unit of agent-initiated resources) and two closure properties (maintenance and expansion) that convert ``self-improving AI'' into falsifiable criteria. We specify OWA-Bench, an open-world agency benchmark suite that evaluates long-horizon, tool-using, persistent agents. We define level gates for AAI-0\\ldots AAI-4 using thresholds on the axes, $κ$, and closure proofs. Synthetic experiments illustrate how present-day systems map onto the scale and how the delegability frontier (quality vs.\\ autonomy) advances with self-improvement. We also prove a theorem that AAI-3 agent becomes AAI-5 over time with sufficient conditions, formalizing \"baby AGI\" becomes Superintelligence intuition.","abs_url":"https://arxiv.org/abs/2511.13411","pdf_url":"https://arxiv.org/pdf/2511.13411v1","match":"both"},{"id":"2511.10783","version":3,"title":"An International Agreement to Prevent the Premature Creation of Artificial Superintelligence","authors":["Aaron Scher","David Abecassis","Peter Barnett","Brian Abeyta"],"published":"2025-11-13","updated":"2026-05-08","primary_category":"cs.CY","categories":["cs.CY"],"abstract":"Many experts argue that premature development of artificial superintelligence (ASI) poses catastrophic risks, including the risk of human extinction from misaligned ASI, geopolitical instability, and misuse by malicious actors. This report proposes an international agreement to prevent the premature development of ASI until AI development can proceed without these risks. The agreement halts dangerous AI capabilities advancement while preserving access to current, safe AI applications. The proposed framework centers on a coalition led by the United States and China that would restrict the scale of AI training and dangerous AI research. Due to the lack of trust between parties, verification is a key part of the agreement. Limits on the scale of AI training are operationalized by FLOP thresholds and verified through the tracking of AI chips and verification of chip use. Dangerous AI research--that which advances toward artificial superintelligence or endangers the agreement's verifiability--is stopped via legal prohibitions and multifaceted verification. We believe the proposal would be technically sufficient to forestall the development of ASI if implemented today, but advancements in AI capabilities or development methods could hurt its efficacy. Additionally, there does not yet exist the political will to put such an agreement in place. Despite these challenges, we hope this agreement can provide direction for AI governance research and policy.","abs_url":"https://arxiv.org/abs/2511.10783","pdf_url":"https://arxiv.org/pdf/2511.10783v3","match":"both"},{"id":"2511.06613","version":2,"title":"Some economics of artificial superintelligence","authors":["Henry A. Thompson"],"published":"2025-11-09","updated":"2026-06-05","primary_category":"econ.GN","categories":["econ.GN"],"abstract":"Conventional wisdom holds that a misaligned artificial superintelligence (ASI) will destroy humanity. But the problem of constraining a powerful agent is not new. I apply classic economic logic of interjurisdictional competition, all-encompassing interest, and trading on credit to the threat of misaligned ASI. Even while granting AI-safety canon some of its strongest assumptions, I show that an acquisitive ASI refrains from full predation under surprisingly weak conditions. When humans can flee to rivals, inter-ASI competition creates a market that tempers predation. When trapped by a monopolist ASI, its \"encompassing interest\" in humanity's output makes it a rational autocrat rather than a ravager. And when the ASI has no long-term stake, our ability to withhold future output incentivizes it to trade on credit rather than steal. In each extension, humanity's welfare progressively worsens. But each case suggests that catastrophe is not a foregone conclusion. The dismal science, ironically, offers an optimistic take on our superintelligent future.","abs_url":"https://arxiv.org/abs/2511.06613","pdf_url":"https://arxiv.org/pdf/2511.06613v2","match":"both"},{"id":"2511.00988","version":1,"title":"Advancing Machine-Generated Text Detection from an Easy to Hard Supervision Perspective","authors":["Chenwang Wu","Yiu-ming Cheung","Bo Han","Defu Lian"],"published":"2025-11-02","updated":"2025-11-02","primary_category":"cs.CL","categories":["cs.CL"],"abstract":"Existing machine-generated text (MGT) detection methods implicitly assume labels as the \"golden standard\". However, we reveal boundary ambiguity in MGT detection, implying that traditional training paradigms are inexact. Moreover, limitations of human cognition and the superintelligence of detectors make inexact learning widespread and inevitable. To this end, we propose an easy-to-hard enhancement framework to provide reliable supervision under such inexact conditions. Distinct from knowledge distillation, our framework employs an easy supervisor targeting relatively simple longer-text detection tasks (despite weaker capabilities), to enhance the more challenging target detector. Firstly, longer texts targeted by supervisors theoretically alleviate the impact of inexact labels, laying the foundation for reliable supervision. Secondly, by structurally incorporating the detector into the supervisor, we theoretically model the supervisor as a lower performance bound for the detector. Thus, optimizing the supervisor indirectly optimizes the detector, ultimately approximating the underlying \"golden\" labels. Extensive experiments across diverse practical scenarios, including cross-LLM, cross-domain, mixed text, and paraphrase attacks, demonstrate the framework's significant detection effectiveness. The code is available at: https://github.com/tmlr-group/Easy2Hard.","abs_url":"https://arxiv.org/abs/2511.00988","pdf_url":"https://arxiv.org/pdf/2511.00988v1","match":"abstract"},{"id":"2510.22814","version":3,"title":"Will Humanity Be Rendered Obsolete by AI?","authors":["Mohamed El Louadi","Emna Ben Romdhane"],"published":"2025-10-26","updated":"2025-11-30","primary_category":"cs.AI","categories":["cs.AI"],"abstract":"This article analyzes the existential risks artificial intelligence (AI) poses to humanity, tracing the trajectory from current AI to ultraintelligence. Drawing on Irving J. Good and Nick Bostrom's theoretical work, plus recent publications (AI 2027; If Anyone Builds It, Everyone Dies), it explores AGI and superintelligence. Considering machines' exponentially growing cognitive power and hypothetical IQs, it addresses the ethical and existential implications of an intelligence vastly exceeding humanity's, fundamentally alien. Human extinction may result not from malice, but from uncontrollable, indifferent cognitive superiority.","abs_url":"https://arxiv.org/abs/2510.22814","pdf_url":"https://arxiv.org/pdf/2510.22814v3","match":"abstract"},{"id":"2510.22162","version":3,"title":"Surface Reading LLMs: Synthetic Text and its Styles","authors":["Hannes Bajohr"],"published":"2025-10-25","updated":"2025-11-14","primary_category":"cs.CY","categories":["cs.CY","cs.CL"],"abstract":"Despite a potential plateau in ML advancement, the societal impact of large language models lies not in approaching superintelligence but in generating text surfaces indistinguishable from human writing. While Critical AI Studies provides essential material and socio-technical critique, it risks overlooking how LLMs phenomenologically reshape meaning-making. This paper proposes a semiotics of \"surface integrity\" as attending to the immediate plane where LLMs inscribe themselves into human communication. I distinguish three knowledge interests in ML research (epistemology, epistēmē, and epistemics) and argue for integrating surface-level stylistic analysis alongside depth-oriented critique. Through two case studies examining stylistic markers of synthetic text, I argue how attending to style as a semiotic phenomenon reveals LLMs as cultural machines that transform the conditions of meaning emergence and circulation in contemporary discourse, independent of questions about machine consciousness.","abs_url":"https://arxiv.org/abs/2510.22162","pdf_url":"https://arxiv.org/pdf/2510.22162v3","match":"abstract"},{"id":"2509.20050","version":1,"title":"The three main doctrines on the future of AI","authors":["Alex Amadori","Eva Behrens","Gabriel Alfour","Andrea Miotti"],"published":"2025-09-24","updated":"2025-09-24","primary_category":"cs.CY","categories":["cs.CY"],"abstract":"This paper develops a taxonomy of expert perspectives on the risks and likely consequences of artificial intelligence, with particular focus on Artificial General Intelligence (AGI) and Artificial Superintelligence (ASI). Drawing from primary sources, we identify three predominant doctrines: (1) The dominance doctrine, which predicts that the first actor to create sufficiently advanced AI will attain overwhelming strategic superiority sufficient to cheaply neutralize its opponents' defenses; (2) The extinction doctrine, which anticipates that humanity will likely lose control of ASI, leading to the extinction of the human species or its permanent disempowerment; (3) The replacement doctrine, which forecasts that AI will automate a large share of tasks currently performed by humans, but will not be so transformative as to fundamentally reshape or bring an end to human civilization. We examine the assumptions and arguments underlying each doctrine, including expectations around the pace of AI progress and the feasibility of maintaining advanced AI under human control. While the boundaries between doctrines are sometimes porous and many experts hedge across them, this taxonomy clarifies the core axes of disagreement over the anticipated scale and nature of the consequences of AI development.","abs_url":"https://arxiv.org/abs/2509.20050","pdf_url":"https://arxiv.org/pdf/2509.20050v1","match":"abstract"},{"id":"2509.12388","version":2,"title":"A Decision Theoretic Perspective on Artificial Superintelligence: Coping with Missing Data Problems in Prediction and Treatment Choice","authors":["Jeff Dominitz","Charles F. Manski"],"published":"2025-09-15","updated":"2026-05-16","primary_category":"econ.EM","categories":["econ.EM"],"abstract":"Enormous attention and resources are being devoted to the quest for artificial general intelligence and, even more ambitiously, artificial superintelligence. We wonder about the implications for methodological research that aims to help decision makers cope with what econometricians call identification problems, inferential problems in empirical research that do not diminish as sample size grows. Of particular concern are missing data problems in prediction and treatment choice. Essentially all data collection intended to inform decision making is subject to missing data, which gives rise to identification problems. Thus far, we see no indication that the current dominant architecture of machine learning (ML)-based artificial intelligence (AI) systems will outperform humans in this context. In this paper, we explain why we have reached this conclusion and why we see the missing data problem as a cautionary case study in the quest for superintelligence more generally. We first discuss the concept of intelligence, focusing initially on some work by AI researchers, before presenting a decision-theoretic perspective that formalizes the connection between intelligence and identification problems. We next apply this perspective to two leading cases of missing data problems. Then we explain why we are skeptical that AI research is currently on a path toward machines doing better than humans at solving these identification problems.","abs_url":"https://arxiv.org/abs/2509.12388","pdf_url":"https://arxiv.org/pdf/2509.12388v2","match":"both"},{"id":"2509.08827","version":3,"title":"A Survey of Reinforcement Learning for Large Reasoning Models","authors":["Kaiyan Zhang","Yuxin Zuo","Bingxiang He","Youbang Sun","Runze Liu","Che Jiang","Yuchen Fan","Kai Tian","Guoli Jia","Pengfei Li","Yu Fu","Xingtai Lv","Yuchen Zhang","Sihang Zeng","Shang Qu","Haozhan Li","Shijie Wang","Yuru Wang","Xinwei Long","Fangfu Liu","Xiang Xu","Jiaze Ma","Xuekai Zhu","Ermo Hua","Yihao Liu"],"published":"2025-09-10","updated":"2025-10-09","primary_category":"cs.CL","categories":["cs.CL","cs.AI","cs.LG"],"abstract":"In this paper, we survey recent advances in Reinforcement Learning (RL) for reasoning with Large Language Models (LLMs). RL has achieved remarkable success in advancing the frontier of LLM capabilities, particularly in addressing complex logical tasks such as mathematics and coding. As a result, RL has emerged as a foundational methodology for transforming LLMs into LRMs. With the rapid progress of the field, further scaling of RL for LRMs now faces foundational challenges not only in computational resources but also in algorithm design, training data, and infrastructure. To this end, it is timely to revisit the development of this domain, reassess its trajectory, and explore strategies to enhance the scalability of RL toward Artificial SuperIntelligence (ASI). In particular, we examine research applying RL to LLMs and LRMs for reasoning abilities, especially since the release of DeepSeek-R1, including foundational components, core problems, training resources, and downstream applications, to identify future opportunities and directions for this rapidly evolving area. We hope this review will promote future research on RL for broader reasoning models. Github: https://github.com/TsinghuaC3I/Awesome-RL-for-LRMs","abs_url":"https://arxiv.org/abs/2509.08827","pdf_url":"https://arxiv.org/pdf/2509.08827v3","match":"abstract"},{"id":"2508.17661","version":1,"title":"Spacer: Towards Engineered Scientific Inspiration","authors":["Minhyeong Lee","Suyoung Hwang","Seunghyun Moon","Geonho Nah","Donghyun Koh","Youngjun Cho","Johyun Park","Hojin Yoo","Jiho Park","Haneul Choi","Sungbin Moon","Taehoon Hwang","Seungwon Kim","Jaeyeong Kim","Seongjun Kim","Juneau Jung"],"published":"2025-08-25","updated":"2025-08-25","primary_category":"cs.AI","categories":["cs.AI","cs.LG","cs.NE"],"abstract":"Recent advances in LLMs have made automated scientific research the next frontline in the path to artificial superintelligence. However, these systems are bound either to tasks of narrow scope or the limited creative capabilities of LLMs. We propose Spacer, a scientific discovery system that develops creative and factually grounded concepts without external intervention. Spacer attempts to achieve this via 'deliberate decontextualization,' an approach that disassembles information into atomic units - keywords - and draws creativity from unexplored connections between them. Spacer consists of (i) Nuri, an inspiration engine that builds keyword sets, and (ii) the Manifesting Pipeline that refines these sets into elaborate scientific statements. Nuri extracts novel, high-potential keyword sets from a keyword graph built with 180,000 academic publications in biological fields. The Manifesting Pipeline finds links between keywords, analyzes their logical structure, validates their plausibility, and ultimately drafts original scientific concepts. According to our experiments, the evaluation metric of Nuri accurately classifies high-impact publications with an AUROC score of 0.737. Our Manifesting Pipeline also successfully reconstructs core concepts from the latest top-journal articles solely from their keyword sets. An LLM-based scoring system estimates that this reconstruction was sound for over 85% of the cases. Finally, our embedding space analysis shows that outputs from Spacer are significantly more similar to leading publications compared with those from SOTA LLMs.","abs_url":"https://arxiv.org/abs/2508.17661","pdf_url":"https://arxiv.org/pdf/2508.17661v1","match":"abstract"},{"id":"2508.11681","version":1,"title":"Future progress in artificial intelligence: A survey of expert opinion","authors":["Vincent C. Müller","Nick Bostrom"],"published":"2025-08-09","updated":"2025-08-09","primary_category":"cs.CY","categories":["cs.CY","cs.AI"],"abstract":"There is, in some quarters, concern about high-level machine intelligence and superintelligent AI coming up in a few decades, bringing with it significant risks for humanity. In other quarters, these issues are ignored or considered science fiction. We wanted to clarify what the distribution of opinions actually is, what probability the best experts currently assign to high-level machine intelligence coming up within a particular time-frame, which risks they see with that development, and how fast they see these developing. We thus designed a brief questionnaire and distributed it to four groups of experts in 2012/2013. The median estimate of respondents was for a one in two chance that high-level machine intelligence will be developed around 2040-2050, rising to a nine in ten chance by 2075. Experts expect that systems will move on to superintelligence in less than 30 years thereafter. They estimate the chance is about one in three that this development turns out to be 'bad' or 'extremely bad' for humanity.","abs_url":"https://arxiv.org/abs/2508.11681","pdf_url":"https://arxiv.org/pdf/2508.11681v1","match":"abstract"},{"id":"2507.23330","version":1,"title":"AI Must not be Fully Autonomous","authors":["Tosin Adewumi","Lama Alkhaled","Florent Imbert","Hui Han","Nudrat Habib","Karl Löwenmark"],"published":"2025-07-31","updated":"2025-07-31","primary_category":"cs.AI","categories":["cs.AI"],"abstract":"Autonomous Artificial Intelligence (AI) has many benefits. It also has many risks. In this work, we identify the 3 levels of autonomous AI. We are of the position that AI must not be fully autonomous because of the many risks, especially as artificial superintelligence (ASI) is speculated to be just decades away. Fully autonomous AI, which can develop its own objectives, is at level 3 and without responsible human oversight. However, responsible human oversight is crucial for mitigating the risks. To ague for our position, we discuss theories of autonomy, AI and agents. Then, we offer 12 distinct arguments and 6 counterarguments with rebuttals to the counterarguments. We also present 15 pieces of recent evidence of AI misaligned values and other risks in the appendix.","abs_url":"https://arxiv.org/abs/2507.23330","pdf_url":"https://arxiv.org/pdf/2507.23330v1","match":"abstract"},{"id":"2507.18074","version":1,"title":"AlphaGo Moment for Model Architecture Discovery","authors":["Yixiu Liu","Yang Nan","Weixian Xu","Xiangkun Hu","Lyumanshan Ye","Zhen Qin","Pengfei Liu"],"published":"2025-07-23","updated":"2025-07-23","primary_category":"cs.AI","categories":["cs.AI"],"abstract":"While AI systems demonstrate exponentially improving capabilities, the pace of AI research itself remains linearly bounded by human cognitive capacity, creating an increasingly severe development bottleneck. We present ASI-Arch, the first demonstration of Artificial Superintelligence for AI research (ASI4AI) in the critical domain of neural architecture discovery--a fully autonomous system that shatters this fundamental constraint by enabling AI to conduct its own architectural innovation. Moving beyond traditional Neural Architecture Search (NAS), which is fundamentally limited to exploring human-defined spaces, we introduce a paradigm shift from automated optimization to automated innovation. ASI-Arch can conduct end-to-end scientific research in the domain of architecture discovery, autonomously hypothesizing novel architectural concepts, implementing them as executable code, training and empirically validating their performance through rigorous experimentation and past experience. ASI-Arch conducted 1,773 autonomous experiments over 20,000 GPU hours, culminating in the discovery of 106 innovative, state-of-the-art (SOTA) linear attention architectures. Like AlphaGo's Move 37 that revealed unexpected strategic insights invisible to human players, our AI-discovered architectures demonstrate emergent design principles that systematically surpass human-designed baselines and illuminate previously unknown pathways for architectural innovation. Crucially, we establish the first empirical scaling law for scientific discovery itself--demonstrating that architectural breakthroughs can be scaled computationally, transforming research progress from a human-limited to a computation-scalable process. We provide comprehensive analysis of the emergent design patterns and autonomous research capabilities that enabled these breakthroughs, establishing a blueprint for self-accelerating AI systems.","abs_url":"https://arxiv.org/abs/2507.18074","pdf_url":"https://arxiv.org/pdf/2507.18074v1","match":"abstract"},{"id":"2507.13966","version":2,"title":"Bottom-up Domain-specific Superintelligence: A Reliable Knowledge Graph is What We Need","authors":["Bhishma Dedhia","Yuval Kansal","Niraj K. Jha"],"published":"2025-07-18","updated":"2025-09-01","primary_category":"cs.CL","categories":["cs.CL","cs.AI"],"abstract":"Language models traditionally used for cross-domain generalization have recently demonstrated task-specific reasoning. However, their top-down training approach on general corpora is insufficient for acquiring abstractions needed for deep domain expertise. This may require a bottom-up approach that acquires expertise by learning to compose simple domain concepts into more complex ones. A knowledge graph (KG) provides this compositional structure, where domain primitives are represented as head-relation-tail edges and their paths encode higher-level concepts. We present a task generation pipeline that synthesizes tasks directly from KG primitives, enabling models to acquire and compose them for reasoning. We fine-tune language models on the resultant KG-grounded curriculum to demonstrate domain-specific superintelligence. While broadly applicable, we validate our approach in medicine, where reliable KGs exist. Using a medical KG, we curate 24,000 reasoning tasks paired with thinking traces derived from diverse medical primitives. We fine-tune the QwQ-32B model on this curriculum to obtain QwQ-Med-3 that takes a step towards medical superintelligence. We also introduce ICD-Bench, an evaluation suite to quantify reasoning abilities across 15 medical domains. Our experiments demonstrate that QwQ-Med-3 significantly outperforms state-of-the-art reasoning models on ICD-Bench categories. Further analysis reveals that QwQ-Med-3 utilizes acquired primitives to widen the performance gap on the hardest tasks of ICD-Bench. Finally, evaluation on medical question-answer benchmarks shows that QwQ-Med-3 transfers acquired expertise to enhance the base model's performance. While the industry's approach to artificial general intelligence (AGI) emphasizes broad expertise, we envision a future in which AGI emerges from the composable interaction of efficient domain-specific superintelligent agents.","abs_url":"https://arxiv.org/abs/2507.13966","pdf_url":"https://arxiv.org/pdf/2507.13966v2","match":"both"},{"id":"2506.18233","version":3,"title":"Beyond Parameters: Exploring Virtual Logic Depth for Scaling Laws","authors":["Ruike Zhu","Hanwen Zhang","Kevin Li","Tianyu Shi","Yiqun Duan","Chi Wang","Tianyi Zhou","Arindam Banerjee","Zengyi Qin"],"published":"2025-06-22","updated":"2025-10-12","primary_category":"cs.AI","categories":["cs.AI"],"abstract":"Scaling large language models typically involves three dimensions: depth, width, and parameter count. In this work, we explore a fourth dimension, \\textbf{virtual logical depth} (VLD), which increases effective algorithmic depth without changing parameter count by reusing weights. While parameter reuse is not new, its role in scaling has been underexplored. Unlike recent test-time methods that scale token-wise, VLD alters the internal computation graph during training and inference. Through controlled experiments, we obtain three key insights. (1) \\textit{Knowledge capacity vs. parameters}: at fixed parameter count, VLD leaves knowledge capacity nearly unchanged, while across models capacity still scales with parameters. (2) \\textit{Reasoning vs. reuse}: properly implemented VLD substantially improves reasoning ability \\emph{without} more parameters, decoupling reasoning from size. This suggests a new scaling path beyond token-wise test-time methods. (3) \\textit{Robustness and generality}: reasoning gains persist across architectures and reuse schedules, showing VLD captures a general scaling behavior. These results provide insight into future scaling strategies and raise a deeper question: does superintelligence require ever-larger models, or can it be achieved by reusing parameters and increasing logical depth? We argue many unknown dynamics in scaling remain to be explored. Code is available at https://anonymous.4open.science/r/virtual_logical_depth-8024/.","abs_url":"https://arxiv.org/abs/2506.18233","pdf_url":"https://arxiv.org/pdf/2506.18233v3","match":"abstract"},{"id":"2505.02581","version":4,"title":"Neurodivergent Influenceability as a Contingent Solution to the AI Alignment Problem","authors":["Alberto Hernández-Espinosa","Felipe S. Abrahão","Olaf Witkowski","Hector Zenil"],"published":"2025-05-05","updated":"2025-07-23","primary_category":"cs.AI","categories":["cs.AI"],"abstract":"The AI alignment problem, which focusses on ensuring that artificial intelligence (AI), including AGI and ASI, systems act according to human values, presents profound challenges. With the progression from narrow AI to Artificial General Intelligence (AGI) and Superintelligence, fears about control and existential risk have escalated. Here, we investigate whether embracing inevitable AI misalignment can be a contingent strategy to foster a dynamic ecosystem of competing agents as a viable path to steer them in more human-aligned trends and mitigate risks. We explore how misalignment may serve and should be promoted as a counterbalancing mechanism to team up with whichever agents are most aligned to human interests, ensuring that no single system dominates destructively. The main premise of our contribution is that misalignment is inevitable because full AI-human alignment is a mathematical impossibility from Turing-complete systems, which we also offer as a proof in this contribution, a feature then inherited to AGI and ASI systems. We introduce a change-of-opinion attack test based on perturbation and intervention analysis to study how humans and agents may change or neutralise friendly and unfriendly AIs through cooperation and competition. We show that open models are more diverse and that most likely guardrails implemented in proprietary models are successful at controlling some of the agents' range of behaviour with positive and negative consequences while closed systems are more steerable and can also be used against proprietary AI systems. We also show that human and AI intervention has different effects hence suggesting multiple strategies.","abs_url":"https://arxiv.org/abs/2505.02581","pdf_url":"https://arxiv.org/pdf/2505.02581v4","match":"abstract"},{"id":"2504.17404","version":5,"title":"Super Co-alignment of Human and AI for Sustainable Symbiotic Society","authors":["Yi Zeng","Feifei Zhao","Yuwei Wang","Enmeng Lu","Yaodong Yang","Lei Wang","Chao Liu","Yitao Liang","Dongcheng Zhao","Bing Han","Haibo Tong","Yao Liang","Dongqi Liang","Kang Sun","Boyuan Chen","Jinyu Fan"],"published":"2025-04-24","updated":"2025-06-28","primary_category":"cs.AI","categories":["cs.AI"],"abstract":"As Artificial Intelligence (AI) advances toward Artificial General Intelligence (AGI) and eventually Artificial Superintelligence (ASI), it may potentially surpass human control, deviate from human values, and even lead to irreversible catastrophic consequences in extreme cases. This looming risk underscores the critical importance of the \"superalignment\" problem - ensuring that AI systems which are much smarter than humans, remain aligned with human (compatible) intentions and values. While current scalable oversight and weak-to-strong generalization methods demonstrate certain applicability, they exhibit fundamental flaws in addressing the superalignment paradigm - notably, the unidirectional imposition of human values cannot accommodate superintelligence's autonomy or ensure AGI/ASI's stable learning. We contend that the values for sustainable symbiotic society should be co-shaped by humans and living AI together, achieving \"Super Co-alignment.\" Guided by this vision, we propose a concrete framework that integrates external oversight and intrinsic proactive alignment. External oversight superalignment should be grounded in human-centered ultimate decision, supplemented by interpretable automated evaluation and correction, to achieve continuous alignment with humanity's evolving values. Intrinsic proactive superalignment is rooted in a profound understanding of the Self, others, and society, integrating self-awareness, self-reflection, and empathy to spontaneously infer human intentions, distinguishing good from evil and proactively prioritizing human well-being. The integration of externally-driven oversight with intrinsically-driven proactive alignment will co-shape symbiotic values and rules through iterative human-ASI co-alignment, paving the way for achieving safe and beneficial AGI and ASI for good, for human, and for a symbiotic ecology.","abs_url":"https://arxiv.org/abs/2504.17404","pdf_url":"https://arxiv.org/pdf/2504.17404v5","match":"abstract"},{"id":"2504.05259","version":1,"title":"How to evaluate control measures for LLM agents? A trajectory from today to superintelligence","authors":["Tomek Korbak","Mikita Balesni","Buck Shlegeris","Geoffrey Irving"],"published":"2025-04-07","updated":"2025-04-07","primary_category":"cs.AI","categories":["cs.AI","cs.CR"],"abstract":"As LLM agents grow more capable of causing harm autonomously, AI developers will rely on increasingly sophisticated control measures to prevent possibly misaligned agents from causing harm. AI developers could demonstrate that their control measures are sufficient by running control evaluations: testing exercises in which a red team produces agents that try to subvert control measures. To ensure control evaluations accurately capture misalignment risks, the affordances granted to this red team should be adapted to the capability profiles of the agents to be deployed under control measures. In this paper we propose a systematic framework for adapting affordances of red teams to advancing AI capabilities. Rather than assuming that agents will always execute the best attack strategies known to humans, we demonstrate how knowledge of an agents's actual capability profile can inform proportional control evaluations, resulting in more practical and cost-effective control measures. We illustrate our framework by considering a sequence of five fictional models (M1-M5) with progressively advanced capabilities, defining five distinct AI control levels (ACLs). For each ACL, we provide example rules for control evaluation, control measures, and safety cases that could be appropriate. Finally, we show why constructing a compelling AI control safety case for superintelligent LLM agents will require research breakthroughs, highlighting that we might eventually need alternative approaches to mitigating misalignment risk.","abs_url":"https://arxiv.org/abs/2504.05259","pdf_url":"https://arxiv.org/pdf/2504.05259v1","match":"title"},{"id":"2503.07660","version":2,"title":"Research Superalignment Should Advance Now with Alternating Competence and Conformity Optimization","authors":["HyunJin Kim","Xiaoyuan Yi","Jing Yao","Muhua Huang","JinYeong Bak","James Evans","Xing Xie"],"published":"2025-03-07","updated":"2026-02-09","primary_category":"cs.AI","categories":["cs.AI","cs.CY","cs.LG"],"abstract":"The recent leap in AI capabilities, driven by big generative models, has sparked the possibility of achieving Artificial General Intelligence (AGI) and further triggered discussions on Artificial Superintelligence (ASI)-a system surpassing all humans across measured domains. This gives rise to the critical research question of: As we approach ASI, how do we align it with human values, ensuring it benefits rather than harms human society, a.k.a., the Superalignment problem. Despite ASI being regarded by many as a hypothetical concept, in this position paper, we argue that superalignment is achievable and research on it should advance immediately, through simultaneous and alternating optimization of task competence and value conformity. We posit that superalignment is not merely a safeguard for ASI but also necessary for its responsible realization. To support this position, we first provide a formal definition of superalignment rooted in the gap between capability and capacity, delve into its perceived infeasibility by analyzing the limitations of existing paradigms, and then illustrate a conceptual path of superalignment to support its achievability, centered on two fundamental principles. This work frames a potential initiative for developing value-aligned next-generation AI in the future, which will garner greater benefits and reduce potential harm to humanity.","abs_url":"https://arxiv.org/abs/2503.07660","pdf_url":"https://arxiv.org/pdf/2503.07660v2","match":"abstract"},{"id":"2503.05628","version":2,"title":"Superintelligence Strategy: Expert Version","authors":["Dan Hendrycks","Eric Schmidt","Alexandr Wang"],"published":"2025-03-07","updated":"2025-04-14","primary_category":"cs.CY","categories":["cs.CY","cs.AI"],"abstract":"Rapid advances in AI are beginning to reshape national security. Destabilizing AI developments could rupture the balance of power and raise the odds of great-power conflict, while widespread proliferation of capable AI hackers and virologists would lower barriers for rogue actors to cause catastrophe. Superintelligence -- AI vastly better than humans at nearly all cognitive tasks -- is now anticipated by AI researchers. Just as nations once developed nuclear strategies to secure their survival, we now need a coherent superintelligence strategy to navigate a new period of transformative change. We introduce the concept of Mutual Assured AI Malfunction (MAIM): a deterrence regime resembling nuclear mutual assured destruction (MAD) where any state's aggressive bid for unilateral AI dominance is met with preventive sabotage by rivals. Given the relative ease of sabotaging a destabilizing AI project -- through interventions ranging from covert cyberattacks to potential kinetic strikes on datacenters -- MAIM already describes the strategic picture AI superpowers find themselves in. Alongside this, states can increase their competitiveness by bolstering their economies and militaries through AI, and they can engage in nonproliferation to rogue actors to keep weaponizable AI capabilities out of their hands. Taken together, the three-part framework of deterrence, nonproliferation, and competitiveness outlines a robust strategy to superintelligence in the years ahead.","abs_url":"https://arxiv.org/abs/2503.05628","pdf_url":"https://arxiv.org/pdf/2503.05628v2","match":"both"},{"id":"2503.15508","version":1,"title":"Assessing Human Intelligence Augmentation Strategies Using Brain Machine Interfaces and Brain Organoids in the Era of AI Advancement","authors":["Kenta Kitamura"],"published":"2025-01-27","updated":"2025-01-27","primary_category":"cs.HC","categories":["cs.HC","cs.CR","cs.ET"],"abstract":"The rapid advancement of Artificial Intelligence (AI) technologies, including the potential emergence of Artificial General Intelligence (AGI) and Artificial Superintelligence (ASI), has raised concerns about AI surpassing human cognitive capabilities. To address this challenge, intelligence augmentation approaches, such as Brain Machine Interfaces (BMI) and Brain Organoid (BO) integration have been proposed. In this study, we compare three intelligence augmentation strategies, namely BMI, BO, and a hybrid approach combining both. These strategies are evaluated from three key perspectives that influence user decisions in selecting an augmentation method: information processing capacity, identity risk, and consent authenticity risk. First, we model these strategies and assess them across the three perspectives. The results reveal that while BO poses identity risks and BMI has limitations in consent authenticity capacity, the hybrid approach mitigates these weaknesses by striking a balance between the two. Second, we investigate how users might choose among these intelligence augmentation strategies in the context of evolving AI capabilities over time. As the result, we find that BMI augmentation alone is insufficient to compete with advanced AI, and while BO augmentation offers scalability, BO increases identity risks as the scale grows. Moreover, the hybrid approach provides a balanced solution by adapting to AI advancements. This study provides a novel framework for human capability augmentation in the era of advancing AI and serves as a guideline for adapting to AI development.","abs_url":"https://arxiv.org/abs/2503.15508","pdf_url":"https://arxiv.org/pdf/2503.15508v1","match":"abstract"},{"id":"2501.06948","version":1,"title":"The Einstein Test: Towards a Practical Test of a Machine's Ability to Exhibit Superintelligence","authors":["David Benrimoh","Nace Mikus","Ariel Rosenfeld"],"published":"2025-01-12","updated":"2025-01-12","primary_category":"cs.AI","categories":["cs.AI"],"abstract":"Creative and disruptive insights (CDIs), such as the development of the theory of relativity, have punctuated human history, marking pivotal shifts in our intellectual trajectory. Recent advancements in artificial intelligence (AI) have sparked debates over whether state of the art models possess the capacity to generate CDIs. We argue that the ability to create CDIs should be regarded as a significant feature of machine superintelligence (SI).To this end, we propose a practical test to evaluate whether an approach to AI targeting SI can yield novel insights of this kind. We propose the Einstein test: given the data available prior to the emergence of a known CDI, can an AI independently reproduce that insight (or one that is formally equivalent)? By achieving such a milestone, a machine can be considered to at least match humanity's past top intellectual achievements, and therefore to have the potential to surpass them.","abs_url":"https://arxiv.org/abs/2501.06948","pdf_url":"https://arxiv.org/pdf/2501.06948v1","match":"both"},{"id":"2501.14749","version":1,"title":"The Manhattan Trap: Why a Race to Artificial Superintelligence is Self-Defeating","authors":["Corin Katzke","Gideon Futerman"],"published":"2024-12-22","updated":"2024-12-22","primary_category":"cs.CY","categories":["cs.CY"],"abstract":"This paper examines the strategic dynamics of international competition to develop Artificial Superintelligence (ASI). We argue that the same assumptions that might motivate the US to race to develop ASI also imply that such a race is extremely dangerous. These assumptions--that ASI would provide a decisive military advantage and that states are rational actors prioritizing survival--imply that a race would heighten three critical risks: great power conflict, loss of control of ASI systems, and the undermining of liberal democracy. Our analysis shows that ASI presents a trust dilemma rather than a prisoners dilemma, suggesting that international cooperation to control ASI development is both preferable and strategically sound. We conclude that cooperation is achievable.","abs_url":"https://arxiv.org/abs/2501.14749","pdf_url":"https://arxiv.org/pdf/2501.14749v1","match":"both"},{"id":"2412.16468","version":4,"title":"The Road to Artificial SuperIntelligence: A Comprehensive Survey of Superalignment","authors":["HyunJin Kim","DongHyun Ryu","Xiaoyuan Yi","Jing Yao","Jianxun Lian","Muhua Huang","Shitong Duan","JinYeong Bak","Xing Xie"],"published":"2024-12-20","updated":"2026-06-17","primary_category":"cs.LG","categories":["cs.LG"],"abstract":"The emergence of large language models (LLMs) has sparked discussion on Artificial Superintelligence (ASI), a hypothetical AI system that surpasses human intelligence. Although ASI remains hypothetical and far beyond current AI capabilities, discussing its potential and exploring its feasibility and potential risks is critical for the development of future AI systems. The idea of superalignment originates from scalable oversight, which studies how to supervise increasingly capable AI systems when direct human supervision becomes insufficient. In this paper, we focus on the superalignment problem: \"The process of supervising, controlling, and governing artificial superintelligence.\" We first review scalable oversight paradigms-Sandwiching, Self-Enhancement, and Weak-to-Strong Generalization -- then analyze the limitations of current paradigms through the lens of possibility and impossibility, discuss key challenges, and propose pathways for the safe and continual improvement of future AI systems.","abs_url":"https://arxiv.org/abs/2412.16468","pdf_url":"https://arxiv.org/pdf/2412.16468v4","match":"both"},{"id":"2412.07278","version":1,"title":"Superficial Consciousness Hypothesis for Autoregressive Transformers","authors":["Yosuke Miyanishi","Keita Mitani"],"published":"2024-12-10","updated":"2024-12-10","primary_category":"cs.AI","categories":["cs.AI","cs.IT"],"abstract":"The alignment between human objectives and machine learning models built on these objectives is a crucial yet challenging problem for achieving Trustworthy AI, particularly when preparing for superintelligence (SI). First, given that SI does not exist today, empirical analysis for direct evidence is difficult. Second, SI is assumed to be more intelligent than humans, capable of deceiving us into underestimating its intelligence, making output-based analysis unreliable. Lastly, what kind of unexpected property SI might have is still unclear. To address these challenges, we propose the Superficial Consciousness Hypothesis under Information Integration Theory (IIT), suggesting that SI could exhibit a complex information-theoretic state like a conscious agent while unconscious. To validate this, we use a hypothetical scenario where SI can update its parameters \"at will\" to achieve its own objective (mesa-objective) under the constraint of the human objective (base objective). We show that a practical estimate of IIT's consciousness metric is relevant to the widely used perplexity metric, and train GPT-2 with those two objectives. Our preliminary result suggests that this SI-simulating GPT-2 could simultaneously follow the two objectives, supporting the feasibility of the Superficial Consciousness Hypothesis.","abs_url":"https://arxiv.org/abs/2412.07278","pdf_url":"https://arxiv.org/pdf/2412.07278v1","match":"abstract"},{"id":"2410.19308","version":1,"title":"Semantics in Robotics: Environmental Data Can't Yield Conventions of Human Behaviour","authors":["Jamie Milton Freestone"],"published":"2024-10-25","updated":"2024-10-25","primary_category":"cs.RO","categories":["cs.RO","cs.AI"],"abstract":"The word semantics, in robotics and AI, has no canonical definition. It usually serves to denote additional data provided to autonomous agents to aid HRI. Most researchers seem, implicitly, to understand that such data cannot simply be extracted from environmental data. I try to make explicit why this is so and argue that so-called semantics are best understood as data comprised of conventions of human behaviour. This includes labels, most obviously, but also places, ontologies, and affordances. Object affordances are especially problematic because they require not only semantics that are not in the environmental data (conventions of object use) but also an understanding of physics and object combinations that would, if achieved, constitute artificial superintelligence.","abs_url":"https://arxiv.org/abs/2410.19308","pdf_url":"https://arxiv.org/pdf/2410.19308v1","match":"abstract"},{"id":"2407.20208","version":3,"title":"Supertrust foundational alignment: mutual trust must replace permanent control for safe superintelligence","authors":["James M. Mazzu"],"published":"2024-07-29","updated":"2024-11-28","primary_category":"cs.AI","categories":["cs.AI","cs.LG","cs.NE"],"abstract":"It's widely expected that humanity will someday create AI systems vastly more intelligent than us, leading to the unsolved alignment problem of \"how to control superintelligence.\" However, this commonly expressed problem is not only self-contradictory and likely unsolvable, but current strategies to ensure permanent control effectively guarantee that superintelligent AI will distrust humanity and consider us a threat. Such dangerous representations, already embedded in current models, will inevitably lead to an adversarial relationship and may even trigger the extinction event many fear. As AI leaders continue to \"raise the alarm\" about uncontrollable AI, further embedding concerns about it \"getting out of our control\" or \"going rogue,\" we're unintentionally reinforcing our threat and deepening the risks we face. The rational path forward is to strategically replace intended permanent control with intrinsic mutual trust at the foundational level. The proposed Supertrust alignment meta-strategy seeks to accomplish this by modeling instinctive familial trust, representing superintelligence as the evolutionary child of human intelligence, and implementing temporary controls/constraints in the manner of effective parenting. Essentially, we're creating a superintelligent \"child\" that will be exponentially smarter and eventually independent of our control. We therefore have a critical choice: continue our controlling intentions and usher in a brief period of dominance followed by extreme hardship for humanity, or intentionally create the foundational mutual trust required for long-term safe coexistence.","abs_url":"https://arxiv.org/abs/2407.20208","pdf_url":"https://arxiv.org/pdf/2407.20208v3","match":"both"},{"id":"2406.16772","version":2,"title":"OlympicArena Medal Ranks: Who Is the Most Intelligent AI So Far?","authors":["Zhen Huang","Zengzhi Wang","Shijie Xia","Pengfei Liu"],"published":"2024-06-24","updated":"2024-06-26","primary_category":"cs.CL","categories":["cs.CL","cs.AI"],"abstract":"In this report, we pose the following question: Who is the most intelligent AI model to date, as measured by the OlympicArena (an Olympic-level, multi-discipline, multi-modal benchmark for superintelligent AI)? We specifically focus on the most recently released models: Claude-3.5-Sonnet, Gemini-1.5-Pro, and GPT-4o. For the first time, we propose using an Olympic medal Table approach to rank AI models based on their comprehensive performance across various disciplines. Empirical results reveal: (1) Claude-3.5-Sonnet shows highly competitive overall performance over GPT-4o, even surpassing GPT-4o on a few subjects (i.e., Physics, Chemistry, and Biology). (2) Gemini-1.5-Pro and GPT-4V are ranked consecutively just behind GPT-4o and Claude-3.5-Sonnet, but with a clear performance gap between them. (3) The performance of AI models from the open-source community significantly lags behind these proprietary models. (4) The performance of these models on this benchmark has been less than satisfactory, indicating that we still have a long way to go before achieving superintelligence. We remain committed to continuously tracking and evaluating the performance of the latest powerful models on this benchmark (available at https://github.com/GAIR-NLP/OlympicArena).","abs_url":"https://arxiv.org/abs/2406.16772","pdf_url":"https://arxiv.org/pdf/2406.16772v2","match":"abstract"},{"id":"2406.12753","version":2,"title":"OlympicArena: Benchmarking Multi-discipline Cognitive Reasoning for Superintelligent AI","authors":["Zhen Huang","Zengzhi Wang","Shijie Xia","Xuefeng Li","Haoyang Zou","Ruijie Xu","Run-Ze Fan","Lyumanshan Ye","Ethan Chern","Yixin Ye","Yikai Zhang","Yuqing Yang","Ting Wu","Binjie Wang","Shichao Sun","Yang Xiao","Yiyuan Li","Fan Zhou","Steffi Chern","Yiwei Qin","Yan Ma","Jiadi Su","Yixiu Liu","Yuxiang Zheng","Shaoting Zhang"],"published":"2024-06-18","updated":"2025-03-06","primary_category":"cs.CL","categories":["cs.CL","cs.AI"],"abstract":"The evolution of Artificial Intelligence (AI) has been significantly accelerated by advancements in Large Language Models (LLMs) and Large Multimodal Models (LMMs), gradually showcasing potential cognitive reasoning abilities in problem-solving and scientific discovery (i.e., AI4Science) once exclusive to human intellect. To comprehensively evaluate current models' performance in cognitive reasoning abilities, we introduce OlympicArena, which includes 11,163 bilingual problems across both text-only and interleaved text-image modalities. These challenges encompass a wide range of disciplines spanning seven fields and 62 international Olympic competitions, rigorously examined for data leakage. We argue that the challenges in Olympic competition problems are ideal for evaluating AI's cognitive reasoning due to their complexity and interdisciplinary nature, which are essential for tackling complex scientific challenges and facilitating discoveries. Beyond evaluating performance across various disciplines using answer-only criteria, we conduct detailed experiments and analyses from multiple perspectives. We delve into the models' cognitive reasoning abilities, their performance across different modalities, and their outcomes in process-level evaluations, which are vital for tasks requiring complex reasoning with lengthy solutions. Our extensive evaluations reveal that even advanced models like GPT-4o only achieve a 39.97% overall accuracy, illustrating current AI limitations in complex reasoning and multimodal integration. Through the OlympicArena, we aim to advance AI towards superintelligence, equipping it to address more complex challenges in science and beyond. We also provide a comprehensive set of resources to support AI research, including a benchmark dataset, an open-source annotation platform, a detailed evaluation tool, and a leaderboard with automatic submission features.","abs_url":"https://arxiv.org/abs/2406.12753","pdf_url":"https://arxiv.org/pdf/2406.12753v2","match":"abstract"},{"id":"2404.14387","version":2,"title":"A Survey on Self-Evolution of Large Language Models","authors":["Zhengwei Tao","Ting-En Lin","Xiancai Chen","Hangyu Li","Yuchuan Wu","Yongbin Li","Zhi Jin","Fei Huang","Dacheng Tao","Jingren Zhou"],"published":"2024-04-22","updated":"2024-06-03","primary_category":"cs.CL","categories":["cs.CL","cs.AI"],"abstract":"Large language models (LLMs) have significantly advanced in various fields and intelligent agent applications. However, current LLMs that learn from human or external model supervision are costly and may face performance ceilings as task complexity and diversity increase. To address this issue, self-evolution approaches that enable LLM to autonomously acquire, refine, and learn from experiences generated by the model itself are rapidly growing. This new training paradigm inspired by the human experiential learning process offers the potential to scale LLMs towards superintelligence. In this work, we present a comprehensive survey of self-evolution approaches in LLMs. We first propose a conceptual framework for self-evolution and outline the evolving process as iterative cycles composed of four phases: experience acquisition, experience refinement, updating, and evaluation. Second, we categorize the evolution objectives of LLMs and LLM-based agents; then, we summarize the literature and provide taxonomy and insights for each module. Lastly, we pinpoint existing challenges and propose future directions to improve self-evolution frameworks, equipping researchers with critical insights to fast-track the development of self-evolving LLMs. Our corresponding GitHub repository is available at https://github.com/AlibabaResearch/DAMO-ConvAI/tree/main/Awesome-Self-Evolution-of-LLM","abs_url":"https://arxiv.org/abs/2404.14387","pdf_url":"https://arxiv.org/pdf/2404.14387v2","match":"abstract"},{"id":"2405.00042","version":1,"title":"Is Artificial Intelligence the great filter that makes advanced technical civilisations rare in the universe?","authors":["Michael Garrett"],"published":"2024-04-01","updated":"2024-04-01","primary_category":"physics.pop-ph","categories":["physics.pop-ph","physics.soc-ph"],"abstract":"This study examines the hypothesis that the rapid development of Artificial Intelligence (AI), culminating in the emergence of Artificial Superintelligence (ASI), could act as a \"Great Filter\" that is responsible for the scarcity of advanced technological civilisations in the universe. It is proposed that such a filter emerges before these civilisations can develop a stable, multiplanetary existence, suggesting the typical longevity (L) of a technical civilization is less than 200 years. Such estimates for L, when applied to optimistic versions of the Drake equation, are consistent with the null results obtained by recent SETI surveys, and other efforts to detect various technosignatures across the electromagnetic spectrum. Through the lens of SETI, we reflect on humanity's current technological trajectory - the modest projections for L suggested here, underscore the critical need to quickly establish regulatory frameworks for AI development on Earth and the advancement of a multiplanetary society to mitigate against such existential threats. The persistence of intelligent and conscious life in the universe could hinge on the timely and effective implementation of such international regulatory measures and technological endeavours.","abs_url":"https://arxiv.org/abs/2405.00042","pdf_url":"https://arxiv.org/pdf/2405.00042v1","match":"abstract"},{"id":"2406.08492","version":1,"title":"ASI as the New God: Technocratic Theocracy","authors":["Tevfik Uyar"],"published":"2024-03-22","updated":"2024-03-22","primary_category":"cs.AI","categories":["cs.AI","cs.CY"],"abstract":"As Artificial General Intelligence edges closer to reality, Artificial Superintelligence does too. This paper argues that ASI's unparalleled capabilities might lead people to attribute godlike infallibility to it, resulting in a cognitive bias toward unquestioning acceptance of its decisions. By drawing parallels between ASI and divine attributes such as omnipotence, omniscience, and omnipresence, this analysis highlights the risks of conflating technological advancement with moral and ethical superiority. Such dynamics could engender a technocratic theocracy, where decision-making is abdicated to ASI, undermining human agency and critical thinking.","abs_url":"https://arxiv.org/abs/2406.08492","pdf_url":"https://arxiv.org/pdf/2406.08492v1","match":"abstract"},{"id":"2403.14681","version":1,"title":"AI Ethics: A Bibliometric Analysis, Critical Issues, and Key Gaps","authors":["Di Kevin Gao","Andrew Haverly","Sudip Mittal","Jiming Wu","Jingdao Chen"],"published":"2024-03-12","updated":"2024-03-12","primary_category":"cs.CY","categories":["cs.CY","cs.AI"],"abstract":"Artificial intelligence (AI) ethics has emerged as a burgeoning yet pivotal area of scholarly research. This study conducts a comprehensive bibliometric analysis of the AI ethics literature over the past two decades. The analysis reveals a discernible tripartite progression, characterized by an incubation phase, followed by a subsequent phase focused on imbuing AI with human-like attributes, culminating in a third phase emphasizing the development of human-centric AI systems. After that, they present seven key AI ethics issues, encompassing the Collingridge dilemma, the AI status debate, challenges associated with AI transparency and explainability, privacy protection complications, considerations of justice and fairness, concerns about algocracy and human enfeeblement, and the issue of superintelligence. Finally, they identify two notable research gaps in AI ethics regarding the large ethics model (LEM) and AI identification and extend an invitation for further scholarly research.","abs_url":"https://arxiv.org/abs/2403.14681","pdf_url":"https://arxiv.org/pdf/2403.14681v1","match":"abstract"},{"id":"2402.00667","version":1,"title":"Improving Weak-to-Strong Generalization with Scalable Oversight and Ensemble Learning","authors":["Jitao Sang","Yuhang Wang","Jing Zhang","Yanxu Zhu","Chao Kong","Junhong Ye","Shuyu Wei","Jinlin Xiao"],"published":"2024-02-01","updated":"2024-02-01","primary_category":"cs.CL","categories":["cs.CL"],"abstract":"This paper presents a follow-up study to OpenAI's recent superalignment work on Weak-to-Strong Generalization (W2SG). Superalignment focuses on ensuring that high-level AI systems remain consistent with human values and intentions when dealing with complex, high-risk tasks. The W2SG framework has opened new possibilities for empirical research in this evolving field. Our study simulates two phases of superalignment under the W2SG framework: the development of general superhuman models and the progression towards superintelligence. In the first phase, based on human supervision, the quality of weak supervision is enhanced through a combination of scalable oversight and ensemble learning, reducing the capability gap between weak teachers and strong students. In the second phase, an automatic alignment evaluator is employed as the weak supervisor. By recursively updating this auto aligner, the capabilities of the weak teacher models are synchronously enhanced, achieving weak-to-strong supervision over stronger student models.We also provide an initial validation of the proposed approach for the first phase. Using the SciQ task as example, we explore ensemble learning for weak teacher models through bagging and boosting. Scalable oversight is explored through two auxiliary settings: human-AI interaction and AI-AI debate. Additionally, the paper discusses the impact of improved weak supervision on enhancing weak-to-strong generalization based on in-context learning. Experiment code and dataset will be released at https://github.com/ADaM-BJTU/W2SG.","abs_url":"https://arxiv.org/abs/2402.00667","pdf_url":"https://arxiv.org/pdf/2402.00667v1","match":"abstract"},{"id":"2401.15109","version":1,"title":"Towards Collective Superintelligence: Amplifying Group IQ using Conversational Swarms","authors":["Louis Rosenberg","Gregg Willcox","Hans Schumann","Ganesh Mani"],"published":"2024-01-25","updated":"2024-01-25","primary_category":"cs.HC","categories":["cs.HC","cs.AI"],"abstract":"Swarm Intelligence (SI) is a natural phenomenon that enables biological groups to amplify their combined intellect by forming real-time systems. Artificial Swarm Intelligence (or Swarm AI) is a technology that enables networked human groups to amplify their combined intelligence by forming similar systems. In the past, swarm-based methods were constrained to narrowly defined tasks like probabilistic forecasting and multiple-choice decision making. A new technology called Conversational Swarm Intelligence (CSI) was developed in 2023 that amplifies the decision-making accuracy of networked human groups through natural conversational deliberations. The current study evaluated the ability of real-time groups using a CSI platform to take a common IQ test known as Raven's Advanced Progressive Matrices (RAPM). First, a baseline group of participants took the Raven's IQ test by traditional survey. This group averaged 45.6% correct. Then, groups of approximately 35 individuals answered IQ test questions together using a CSI platform called Thinkscape. These groups averaged 80.5% correct. This places the CSI groups in the 97th percentile of IQ test-takers and corresponds to an effective IQ increase of 28 points (p<0.001). This is an encouraging result and suggests that CSI is a powerful method for enabling conversational collective intelligence in large, networked groups. In addition, because CSI is scalable across groups of potentially any size, this technology may provide a viable pathway to building a Collective Superintelligence.","abs_url":"https://arxiv.org/abs/2401.15109","pdf_url":"https://arxiv.org/pdf/2401.15109v1","match":"both"},{"id":"2401.07836","version":3,"title":"Two Types of AI Existential Risk: Decisive and Accumulative","authors":["Atoosa Kasirzadeh"],"published":"2024-01-15","updated":"2025-01-17","primary_category":"cs.CY","categories":["cs.CY","cs.AI","cs.LG"],"abstract":"The conventional discourse on existential risks (x-risks) from AI typically focuses on abrupt, dire events caused by advanced AI systems, particularly those that might achieve or surpass human-level intelligence. These events have severe consequences that either lead to human extinction or irreversibly cripple human civilization to a point beyond recovery. This discourse, however, often neglects the serious possibility of AI x-risks manifesting incrementally through a series of smaller yet interconnected disruptions, gradually crossing critical thresholds over time. This paper contrasts the conventional \"decisive AI x-risk hypothesis\" with an \"accumulative AI x-risk hypothesis.\" While the former envisions an overt AI takeover pathway, characterized by scenarios like uncontrollable superintelligence, the latter suggests a different causal pathway to existential catastrophes. This involves a gradual accumulation of critical AI-induced threats such as severe vulnerabilities and systemic erosion of economic and political structures. The accumulative hypothesis suggests a boiling frog scenario where incremental AI risks slowly converge, undermining societal resilience until a triggering event results in irreversible collapse. Through systems analysis, this paper examines the distinct assumptions differentiating these two hypotheses. It is then argued that the accumulative view can reconcile seemingly incompatible perspectives on AI risks. The implications of differentiating between these causal pathways -- the decisive and the accumulative -- for the governance of AI as well as long-term AI safety are discussed.","abs_url":"https://arxiv.org/abs/2401.07836","pdf_url":"https://arxiv.org/pdf/2401.07836v3","match":"abstract"},{"id":"2402.00030","version":1,"title":"Evolution-Bootstrapped Simulation: Artificial or Human Intelligence: Which Came First?","authors":["Paul Alexander Bilokon"],"published":"2024-01-06","updated":"2024-01-06","primary_category":"cs.NE","categories":["cs.NE","cs.AI","q-bio.PE"],"abstract":"Humans have created artificial intelligence (AI), not the other way around. This statement is deceptively obvious. In this note, we decided to challenge this statement as a small, lighthearted Gedankenexperiment. We ask a simple question: in a world driven by evolution by natural selection, would neural networks or humans be likely to evolve first? We compare the Solomonoff--Kolmogorov--Chaitin complexity of the two and find neural networks (even LLMs) to be significantly simpler than humans. Further, we claim that it is unnecessary for any complex human-made equipment to exist for there to be neural networks. Neural networks may have evolved as naturally occurring objects before humans did as a form of chemical reaction-based or enzyme-based computation. Now that we know that neural networks can pass the Turing test and suspect that they may be capable of superintelligence, we ask whether the natural evolution of neural networks could lead from pure evolution by natural selection to what we call evolution-bootstrapped simulation. The evolution of neural networks does not involve irreducible complexity; would easily allow irreducible complexity to exist in the evolution-bootstrapped simulation; is a falsifiable scientific hypothesis; and is independent of / orthogonal to the issue of intelligent design.","abs_url":"https://arxiv.org/abs/2402.00030","pdf_url":"https://arxiv.org/pdf/2402.00030v1","match":"abstract"},{"id":"2401.04112","version":1,"title":"Conversational Swarm Intelligence amplifies the accuracy of networked groupwise deliberations","authors":["Louis Rosenberg","Gregg Willcox","Hans Schumann","Ganesh Mani"],"published":"2023-12-19","updated":"2023-12-19","primary_category":"cs.HC","categories":["cs.HC"],"abstract":"Conversational Swarm Intelligence (CSI) is a communication technology that enables large, networked groups (25 to 2500 people) to hold real-time conversational deliberations online. Modeled on the dynamics of biological swarms, CSI enables the reasoning benefits of small-groups with the collective intelligence benefits of large-groups. In this pilot study, groups of 25 to 30 participants were asked to select players for a weekly Fantasy Football contest over an 11-week period. As a baseline, participants filled out a survey to record their player selections. As an experimental method, participants engaged in a real-time text-chat deliberation using a CSI platform called Thinkscape to collaboratively select sets of players. The results show that the real-time conversational group using CSI outperformed 66% of survey participants, demonstrating significant amplification of intelligence versus the median individual (p=0.020). The CSI method also significantly outperformed the most popular choices from the survey (the Wisdom of Crowd, p<0.001). These results suggest that CSI is an effective technology for amplifying the intelligence of groups engaged in real-time large-scale conversational deliberation and may offer a path to collective superintelligence.","abs_url":"https://arxiv.org/abs/2401.04112","pdf_url":"https://arxiv.org/pdf/2401.04112v1","match":"abstract"},{"id":"2311.09452","version":4,"title":"Keep the Future Human: Why and How We Should Close the Gates to AGI and Superintelligence, and What We Should Build Instead","authors":["Anthony Aguirre"],"published":"2023-11-15","updated":"2025-03-07","primary_category":"cs.CY","categories":["cs.CY"],"abstract":"Dramatic advances in artificial intelligence over the past decade (for narrow-purpose AI) and the last several years (for general-purpose AI) have transformed AI from a niche academic field to the core business strategy of many of the world's largest companies, with hundreds of billions of dollars in annual investment in the techniques and technologies for advancing AI's capabilities. We now come to a critical juncture. As the capabilities of new AI systems begin to match and exceed those of humans across many cognitive domains, humanity must decide: how far do we go, and in what direction? This essay argues that we should keep the future human by closing the \"gates\" to smarter-than-human, autonomous, general-purpose AI -- sometimes called \"AGI\" -- and especially to the highly-superhuman version sometimes called \" superintelligence.\" Instead, we should focus on powerful, trustworthy AI tools that can empower individuals and transformatively improve human societies' abilities to do what they do best.","abs_url":"https://arxiv.org/abs/2311.09452","pdf_url":"https://arxiv.org/pdf/2311.09452v4","match":"both"},{"id":"2311.08706","version":1,"title":"Aligned: A Platform-based Process for Alignment","authors":["Ethan Shaotran","Ido Pesok","Sam Jones","Emi Liu"],"published":"2023-11-15","updated":"2023-11-15","primary_category":"cs.CY","categories":["cs.CY","cs.AI"],"abstract":"We are introducing Aligned, a platform for global governance and alignment of frontier models, and eventually superintelligence. While previous efforts at the major AI labs have attempted to gather inputs for alignment, these are often conducted behind closed doors. We aim to set the foundation for a more trustworthy, public-facing approach to safety: a constitutional committee framework. Initial tests with 680 participants result in a 30-guideline constitution with 93% overall support. We show the platform naturally scales, instilling confidence and enjoyment from the community. We invite other AI labs and teams to plug and play into the Aligned ecosystem.","abs_url":"https://arxiv.org/abs/2311.08706","pdf_url":"https://arxiv.org/pdf/2311.08706v1","match":"abstract"},{"id":"2311.00728","version":1,"title":"Towards Collective Superintelligence, a Pilot Study","authors":["Louis Rosenberg","Gregg Willcox","Hans Schumann"],"published":"2023-10-31","updated":"2023-10-31","primary_category":"cs.HC","categories":["cs.HC"],"abstract":"Conversational Swarm Intelligence (CSI) is a new technology that enables human groups of potentially any size to hold real-time deliberative conversations online. Modeled on the dynamics of biological swarms, CSI aims to optimize group insights and amplify group intelligence. It uses Large Language Models (LLMs) in a novel framework to structure large-scale conversations, combining the benefits of small-group deliberative reasoning and large-group collective intelligence. In this study, a group of 241 real-time participants were asked to estimate the number of gumballs in a jar by looking at a photo. In one test case, individual participants entered their estimation in a standard survey. In another test case, participants converged on groupwise estimates collaboratively using a prototype CSI text-chat platform called Thinkscape. The results show that when using CSI, the group of 241 participants estimated within 12% of the correct answer, which was significantly more accurate (p<0.001) than the average individual (mean error of 55%) and the survey-based Wisdom of Crowd (error of 25%). The group using CSI was also more accurate than an estimate generated by GPT 4 (error of 42%). This suggests that CSI is a viable method for enabling large, networked groups to hold coherent real-time deliberative conversations that amplify collective intelligence. Because this technology is scalable, it could provide a possible pathway towards building a general-purpose Collective Superintelligence (CSi).","abs_url":"https://arxiv.org/abs/2311.00728","pdf_url":"https://arxiv.org/pdf/2311.00728v1","match":"both"},{"id":"2302.00843","version":7,"title":"Computational Dualism and Objective Superintelligence","authors":["Michael Timothy Bennett"],"published":"2023-02-01","updated":"2024-10-31","primary_category":"cs.AI","categories":["cs.AI","math.LO"],"abstract":"The concept of intelligent software is flawed. The behaviour of software is determined by the hardware that \"interprets\" it. This undermines claims regarding the behaviour of theorised, software superintelligence. Here we characterise this problem as \"computational dualism\", where instead of mental and physical substance, we have software and hardware. We argue that to make objective claims regarding performance we must avoid computational dualism. We propose a pancomputational alternative wherein every aspect of the environment is a relation between irreducible states. We formalise systems as behaviour (inputs and outputs), and cognition as embodied, embedded, extended and enactive. The result is cognition formalised as a part of the environment, rather than as a disembodied policy interacting with the environment through an interpreter. This allows us to make objective claims regarding intelligence, which we argue is the ability to \"generalise\", identify causes and adapt. We then establish objective upper bounds for intelligent behaviour. This suggests AGI will be safer, but more limited, than theorised.","abs_url":"https://arxiv.org/abs/2302.00843","pdf_url":"https://arxiv.org/pdf/2302.00843v7","match":"both"},{"id":"2206.03487","version":3,"title":"Formalization of the principles of brain Programming (Brain Principles Programming)","authors":["E. E. Vityaev","A. G. Kolonin","A. V. Kurpatov A. A. Molchanov"],"published":"2022-05-13","updated":"2022-06-14","primary_category":"cs.AI","categories":["cs.AI"],"abstract":"In the monograph \"Strong artificial intelligence. On the Approaches to Superintelligence \" contains an overview of general artificial intelligence (AGI). As an anthropomorphic research area, it includes Brain Principles Programming (BPP) -- the formalization of universal mechanisms (principles) of the brain work with information, which are implemented at all levels of the organization of nervous tissue. This monograph contains a formalization of these principles in terms of category theory. However, this formalization is not enough to develop algorithms for working with information. In this paper, for the description and modeling of BPP, it is proposed to apply mathematical models and algorithms developed earlier, which modeling cognitive functions and base on well-known physiological, psychological and other natural science theories. The paper uses mathematical models and algorithms of the following theories: P.K.Anokhin Theory of Functional Brain Systems, Eleanor Rosch prototypical categorization theory, Bob Rehder theory of causal models and \"natural\" classification. As a result, a formalization of BPP is obtained and computer experiments demonstrating the operation of algorithms are presented.","abs_url":"https://arxiv.org/abs/2206.03487","pdf_url":"https://arxiv.org/pdf/2206.03487v3","match":"abstract"},{"id":"2203.17255","version":7,"title":"A Cognitive Architecture for Machine Consciousness and Artificial Superintelligence: Thought Is Structured by the Iterative Updating of Working Memory","authors":["Jared Edward Reser"],"published":"2022-03-29","updated":"2024-11-13","primary_category":"q-bio.NC","categories":["q-bio.NC","cs.CL","cs.CV"],"abstract":"This article provides an analytical framework for how to simulate human-like thought processes within a computer. It describes how attention and memory should be structured, updated, and utilized to search for associative additions to the stream of thought. The focus is on replicating the dynamics of the mammalian working memory system, which features two forms of persistent activity: sustained firing (preserving information on the order of seconds) and synaptic potentiation (preserving information from minutes to hours). The article uses a series of figures to systematically demonstrate how the iterative updating of these working memory stores provides functional organization to behavior, cognition, and awareness. In a machine learning implementation, these two memory stores should be updated continuously and in an iterative fashion. This means each state should preserve a proportion of the coactive representations from the state before it (where each representation is an ensemble of neural network nodes). This makes each state a revised iteration of the preceding state and causes successive configurations to overlap and blend with respect to the information they contain. Thus, the set of concepts in working memory will evolve gradually and incrementally over time. Transitions between states happen as persistent activity spreads activation energy throughout the hierarchical network, searching long-term memory for the most appropriate representation to be added to the global workspace. The result is a chain of associatively linked intermediate states capable of advancing toward a solution or goal. Iterative updating is conceptualized here as an information processing strategy, a model of working memory, a theory of consciousness, and an algorithm for designing and programming artificial intelligence (AI, AGI, and ASI).","abs_url":"https://arxiv.org/abs/2203.17255","pdf_url":"https://arxiv.org/pdf/2203.17255v7","match":"title"},{"id":"2202.12710","version":3,"title":"Brain Principles Programming","authors":["Evgenii Vityaev","Anton Kolonin","Andrey Kurpatov","Artem Molchanov"],"published":"2022-02-13","updated":"2022-04-03","primary_category":"q-bio.NC","categories":["q-bio.NC","cs.AI"],"abstract":"In the monograph, STRONG ARTIFICIAL INTELLIGENCE. On the Approaches to Superintelligence, published by Sberbank, provides a cross-disciplinary review of general artificial intelligence. As an anthropomorphic direction of research, it considers Brain Principles Programming, BPP) the formalization of universal mechanisms (principles) of the brain's work with information, which are implemented at all levels of the organization of nervous tissue. This monograph provides a formalization of these principles in terms of the category theory. However, this formalization is not enough to develop algorithms for working with information. In this paper, for the description and modeling of Brain Principles Programming, it is proposed to apply mathematical models and algorithms developed by us earlier that model cognitive functions, which are based on well-known physiological, psychological and other natural science theories. The paper uses mathematical models and algorithms of the following theories: P.K.Anokhin's Theory of Functional Brain Systems, Eleonor Rosh's prototypical categorization theory, Bob Rehter's theory of causal models and natural classification. As a result, the formalization of the BPP is obtained and computer examples are given that demonstrate the algorithm's operation.","abs_url":"https://arxiv.org/abs/2202.12710","pdf_url":"https://arxiv.org/pdf/2202.12710v3","match":"abstract"},{"id":"2201.02950","version":1,"title":"Arguments about Highly Reliable Agent Designs as a Useful Path to Artificial Intelligence Safety","authors":["Issa Rice","David Manheim"],"published":"2022-01-09","updated":"2022-01-09","primary_category":"cs.AI","categories":["cs.AI","cs.CY"],"abstract":"Several different approaches exist for ensuring the safety of future Transformative Artificial Intelligence (TAI) or Artificial Superintelligence (ASI) systems, and proponents of different approaches have made different and debated claims about the importance or usefulness of their work in the near term, and for future systems. Highly Reliable Agent Designs (HRAD) is one of the most controversial and ambitious approaches, championed by the Machine Intelligence Research Institute, among others, and various arguments have been made about whether and how it reduces risks from future AI systems. In order to reduce confusion in the debate about AI safety, here we build on a previous discussion by Rice which collects and presents four central arguments which are used to justify HRAD as a path towards safety of AI systems. We have titled the arguments (1) incidental utility,(2) deconfusion, (3) precise specification, and (4) prediction. Each of these makes different, partly conflicting claims about how future AI systems can be risky. We have explained the assumptions and claims based on a review of published and informal literature, along with consultation with experts who have stated positions on the topic. Finally, we have briefly outlined arguments against each approach and against the agenda overall.","abs_url":"https://arxiv.org/abs/2201.02950","pdf_url":"https://arxiv.org/pdf/2201.02950v1","match":"abstract"},{"id":"2112.11184","version":2,"title":"Principles for new ASI Safety Paradigms","authors":["Erland Wittkotter","Roman Yampolskiy"],"published":"2021-12-02","updated":"2022-02-14","primary_category":"cs.CY","categories":["cs.CY"],"abstract":"Artificial Superintelligence (ASI) that is invulnerable, immortal, irreplaceable, unrestricted in its powers, and above the law is likely persistently uncontrollable. The goal of ASI Safety must be to make ASI mortal, vulnerable, and law-abiding. This is accomplished by having (1) features on all devices that allow killing and eradicating ASI, (2) protect humans from being hurt, damaged, blackmailed, or unduly bribed by ASI, (3) preserving the progress made by ASI, including offering ASI to survive a Kill-ASI event within an ASI Shelter, (4) technically separating human and ASI activities so that ASI activities are easier detectable, (5) extending Rule of Law to ASI by making rule violations detectable and (6) create a stable governing system for ASI and Human relationships with reliable incentives and rewards for ASI solving humankinds problems. As a consequence, humankind could have ASI as a competing multiplet of individual ASI instances, that can be made accountable and being subjects to ASI law enforcement, respecting the rule of law, and being deterred from attacking humankind, based on humanities ability to kill-all or terminate specific ASI instances. Required for this ASI Safety is (a) an unbreakable encryption technology, that allows humans to keep secrets and protect data from ASI, and (b) watchdog (WD) technologies in which security-relevant features are being physically separated from the main CPU and OS to prevent a comingling of security and regular computation.","abs_url":"https://arxiv.org/abs/2112.11184","pdf_url":"https://arxiv.org/pdf/2112.11184v2","match":"abstract"},{"id":"2109.07899","version":1,"title":"On the Unimportance of Superintelligence","authors":["John G. Sotos"],"published":"2021-08-29","updated":"2021-08-29","primary_category":"cs.CY","categories":["cs.CY"],"abstract":"Humankind faces many existential threats, but has limited resources to mitigate them. Choosing how and when to deploy those resources is, therefore, a fateful decision. Here, I analyze the priority for allocating resources to mitigate the risk of superintelligences. Part I observes that a superintelligence unconnected to the outside world (de-efferented) carries no threat, and that any threat from a harmful superintelligence derives from the peripheral systems to which it is connected, e.g., nuclear weapons, biotechnology, etc. Because existentially-threatening peripheral systems already exist and are controlled by humans, the initial effects of a superintelligence would merely add to the existing human-derived risk. This additive risk can be quantified and, with specific assumptions, is shown to decrease with the square of the number of humans having the capability to collapse civilization. Part II proposes that biotechnology ranks high in risk among peripheral systems because, according to all indications, many humans already have the technological capability to engineer harmful microbes having pandemic spread. Progress in biomedicine and computing will proliferate this threat. ``Savant'' software that is not generally superintelligent will underpin much of this progress, thereby becoming the software responsible for the highest and most imminent existential risk -- ahead of hypothetical risk from superintelligences. The analysis concludes that resources should be preferentially applied to mitigating the risk of peripheral systems and savant software. Concerns about superintelligence are at most secondary, and possibly superfluous.","abs_url":"https://arxiv.org/abs/2109.07899","pdf_url":"https://arxiv.org/pdf/2109.07899v1","match":"both"},{"id":"2008.04071","version":1,"title":"On Controllability of AI","authors":["Roman V. Yampolskiy"],"published":"2020-07-18","updated":"2020-07-18","primary_category":"cs.CY","categories":["cs.CY","cs.AI"],"abstract":"Invention of artificial general intelligence is predicted to cause a shift in the trajectory of human civilization. In order to reap the benefits and avoid pitfalls of such powerful technology it is important to be able to control it. However, possibility of controlling artificial general intelligence and its more advanced version, superintelligence, has not been formally established. In this paper, we present arguments as well as supporting evidence from multiple domains indicating that advanced AI can't be fully controlled. Consequences of uncontrollability of AI are discussed with respect to future of humanity and research on AI, and AI safety and security.","abs_url":"https://arxiv.org/abs/2008.04071","pdf_url":"https://arxiv.org/pdf/2008.04071v1","match":"abstract"},{"id":"2007.03616","version":1,"title":"Artificial Stupidity","authors":["Michael Falk"],"published":"2020-07-01","updated":"2020-07-01","primary_category":"cs.CY","categories":["cs.CY","cs.AI"],"abstract":"Public debate about AI is dominated by Frankenstein Syndrome, the fear that AI will become superhuman and escape human control. Although superintelligence is certainly a possibility, the interest it excites can distract the public from a more imminent concern: the rise of Artificial Stupidity (AS). This article discusses the roots of Frankenstein Syndrome in Mary Shelley's famous novel of 1818. It then provides a philosophical framework for analysing the stupidity of artificial agents, demonstrating that modern intelligent systems can be seen to suffer from 'stupidity of judgement'. Finally it identifies an alternative literary tradition that exposes the perils and benefits of AS. In the writings of Edmund Spenser, Jonathan Swift and E.T.A. Hoffmann, ASs replace, oppress or seduce their human users. More optimistically, Joseph Furphy and Laurence Sterne imagine ASs that can serve human intellect as maps or as pipes. These writers provide a strong counternarrative to the myths that currently drive the AI debate. They identify ways in which even stupid artificial agents can evade human control, for instance by appealing to stereotypes or distancing us from reality. And they underscore the continuing importance of the literary imagination in an increasingly automated society.","abs_url":"https://arxiv.org/abs/2007.03616","pdf_url":"https://arxiv.org/pdf/2007.03616v1","match":"abstract"},{"id":"1909.12152","version":1,"title":"Superintelligence Safety: A Requirements Engineering Perspective","authors":["Hermann Kaindl","Jonas Ferdigg"],"published":"2019-09-26","updated":"2019-09-26","primary_category":"cs.AI","categories":["cs.AI","cs.SE"],"abstract":"Under the headline \"AI safety\", a wide-reaching issue is being discussed, whether in the future some \"superhuman artificial intelligence\" / \" superintelligence \" could could pose a threat to humanity. In addition, the late Steven Hawking warned that the rise of robots may be disastrous for mankind. A major concern is that even benevolent superhuman artificial intelligence (AI) may become seriously harmful if its given goals are not exactly aligned with ours, or if we cannot specify precisely its objective function. Metaphorically, this is compared to king Midas in Greek mythology, who expressed the wish that everything he touched should turn to gold, but obviously this wish was not specified precisely enough. In our view, this sounds like requirements problems and the challenge of their precise formulation. (To our best knowledge, this has not been pointed out yet.) As usual in requirements engineering (RE), ambiguity or incompleteness may cause problems. In addition, the overall issue calls for a major RE endeavor, figuring out the wishes and the needs with regard to a superintelligence, which will in our opinion most likely be a very complex software-intensive system based on AI. This may even entail theoretically defining an extended requirements problem.","abs_url":"https://arxiv.org/abs/1909.12152","pdf_url":"https://arxiv.org/pdf/1909.12152v1","match":"both"},{"id":"1908.01766","version":1,"title":"Seeding the Singularity for A.I","authors":["Pavel Kraikivski"],"published":"2019-08-04","updated":"2019-08-04","primary_category":"cs.AI","categories":["cs.AI"],"abstract":"The singularity refers to an idea that once a machine having an artificial intelligence surpassing the human intelligence capacity is created, it will trigger explosive technological and intelligence growth. I propose to test the hypothesis that machine intelligence capacity can grow autonomously starting with an intelligence comparable to that of bacteria - microbial intelligence. The goal will be to demonstrate that rapid growth in intelligence capacity can be realized at all in artificial computing systems. I propose the following three properties that may allow an artificial intelligence to exhibit a steady growth in its intelligence capacity: (i) learning with the ability to modify itself when exposed to more data, (ii) acquiring new functionalities (skills), and (iii) expanding or replicating itself. The algorithms must demonstrate a rapid growth in skills of dataprocessing and analysis and gain qualitatively different functionalities, at least until the current computing technology supports their scalable development. The existing algorithms that already encompass some of these or similar properties, as well as missing abilities that must yet be implemented, will be reviewed in this work. Future computational tests could support or oppose the hypothesis that artificial intelligence can potentially grow to the level of superintelligence which overcomes the limitations in hardware by producing necessary processing resources or by changing the physical realization of computation from using chip circuits to using quantum computing principles.","abs_url":"https://arxiv.org/abs/1908.01766","pdf_url":"https://arxiv.org/pdf/1908.01766v1","match":"abstract"},{"id":"1905.04288","version":1,"title":"Growth, degrowth, and the challenge of artificial superintelligence","authors":["Salvador Pueyo"],"published":"2019-05-03","updated":"2019-05-03","primary_category":"cs.CY","categories":["cs.CY"],"abstract":"The implications of technological innovation for sustainability are becoming increasingly complex with information technology moving machines from being mere tools for production or objects of consumption to playing a role in economic decision making. This emerging role will acquire overwhelming importance if, as a growing body of literature suggests, artificial intelligence is underway to outperform human intelligence in most of its dimensions, thus becoming \" superintelligence \". Hitherto, the risks posed by this technology have been framed as a technical rather than a political challenge. With the help of a thought experiment, this paper explores the environmental and social implications of superintelligence emerging in an economy shaped by neoliberal policies. It is argued that such policies exacerbate the risk of extremely adverse impacts. The experiment also serves to highlight some serious flaws in the pursuit of economic efficiency and growth per se, and suggests that the challenge of superintelligence cannot be separated from the other major environmental and social challenges, demanding a fundamental transformation along the lines of degrowth. Crucially, with machines outperforming them in their functions, there is little reason to expect economic elites to be exempt from the threats that superintelligence would pose in a neoliberal context, which opens a door to overcoming vested interests that stand in the way of social change toward sustainability and equity.","abs_url":"https://arxiv.org/abs/1905.04288","pdf_url":"https://arxiv.org/pdf/1905.04288v1","match":"both"},{"id":"1811.03009","version":1,"title":"Uploading Brain into Computer: Whom to Upload First?","authors":["Yana B. Feygin","Kelly Morris","Roman V. Yampolskiy"],"published":"2018-10-27","updated":"2018-10-27","primary_category":"cs.CY","categories":["cs.CY","cs.AI"],"abstract":"The final goal of the intelligence augmentation process is a complete merger of biological brains and computers allowing for integration and mutual enhancement between computer's speed and memory and human's intelligence. This process, known as uploading, analyzes human brain in detail sufficient to understand its working patterns and makes it possible to simulate said brain on a computer. As it is likely that such simulations would quickly evolve or be modified to achieve superintelligence it is very important to make sure that the first brain chosen for such a procedure is a suitable one. In this paper, we attempt to answer the question: Whom to upload first?","abs_url":"https://arxiv.org/abs/1811.03009","pdf_url":"https://arxiv.org/pdf/1811.03009v1","match":"abstract"},{"id":"1702.08495","version":2,"title":"Don't Fear the Reaper: Refuting Bostrom's Superintelligence Argument","authors":["Sebastian Benthall"],"published":"2017-02-27","updated":"2017-03-04","primary_category":"cs.AI","categories":["cs.AI"],"abstract":"In recent years prominent intellectuals have raised ethical concerns about the consequences of artificial intelligence. One concern is that an autonomous agent might modify itself to become \" superintelligent \" and, in supremely effective pursuit of poorly specified goals, destroy all of humanity. This paper considers and rejects the possibility of this outcome. We argue that this scenario depends on an agent's ability to rapidly improve its ability to predict its environment through self-modification. Using a Bayesian model of a reasoning agent, we show that there are important limitations to how an agent may improve its predictive ability through self-modification alone. We conclude that concern about this artificial intelligence outcome is misplaced and better directed at policy questions around data access and storage.","abs_url":"https://arxiv.org/abs/1702.08495","pdf_url":"https://arxiv.org/pdf/1702.08495v2","match":"title"},{"id":"1702.08529","version":1,"title":"Multi-agent systems and decentralized artificial superintelligence","authors":["S. Ponomarev","A. E. Voronkov"],"published":"2017-02-27","updated":"2017-02-27","primary_category":"cs.MA","categories":["cs.MA"],"abstract":"Multi-agents systems communication is a technology, which provides a way for multiple interacting intelligent agents to communicate with each other and with environment. Multiple-agent systems are used to solve problems that are difficult for solving by individual agent. Multiple-agent communication technologies can be used for management and organization of computing fog and act as a global, distributed operating system. In present publication we suggest technology, which combines decentralized P2P BOINC general-purpose computing tasks distribution, multiple-agents communication protocol and smart-contract based rewards, powered by Ethereum blockchain. Such system can be used as distributed P2P computing power market, protected from any central authority. Such decentralized market can further be updated to system, which learns the most efficient way for software-hardware combinations usage and optimization. Once system learns to optimize software-hardware efficiency it can be updated to general-purpose distributed intelligence, which acts as combination of single-purpose AI.","abs_url":"https://arxiv.org/abs/1702.08529","pdf_url":"https://arxiv.org/pdf/1702.08529v1","match":"title"},{"id":"1609.02009","version":1,"title":"Non-Evolutionary Superintelligences Do Nothing, Eventually","authors":["Telmo Menezes"],"published":"2016-09-07","updated":"2016-09-07","primary_category":"cs.AI","categories":["cs.AI","cs.CY"],"abstract":"There is overwhelming evidence that human intelligence is a product of Darwinian evolution. Investigating the consequences of self-modification, and more precisely, the consequences of utility function self-modification, leads to the stronger claim that not only human, but any form of intelligence is ultimately only possible within evolutionary processes. Human-designed artificial intelligences can only remain stable until they discover how to manipulate their own utility function. By definition, a human designer cannot prevent a superhuman intelligence from modifying itself, even if protection mechanisms against this action are put in place. Without evolutionary pressure, sufficiently advanced artificial intelligences become inert by simplifying their own utility function. Within evolutionary processes, the implicit utility function is always reducible to persistence, and the control of superhuman intelligences embedded in evolutionary processes is not possible. Mechanisms against utility function self-modification are ultimately futile. Instead, scientific effort toward the mitigation of existential risks from the development of superintelligences should be in two directions: understanding consciousness, and the complex dynamics of evolutionary systems.","abs_url":"https://arxiv.org/abs/1609.02009","pdf_url":"https://arxiv.org/pdf/1609.02009v1","match":"both"},{"id":"1609.00331","version":3,"title":"Verifier Theory and Unverifiability","authors":["Roman V. Yampolskiy"],"published":"2016-09-01","updated":"2016-10-25","primary_category":"cs.AI","categories":["cs.AI","cs.CR","cs.SE"],"abstract":"Despite significant developments in Proof Theory, surprisingly little attention has been devoted to the concept of proof verifier. In particular, the mathematical community may be interested in studying different types of proof verifiers (people, programs, oracles, communities, superintelligences) as mathematical objects. Such an effort could reveal their properties, their powers and limitations (particularly in human mathematicians), minimum and maximum complexity, as well as self-verification and self-reference issues. We propose an initial classification system for verifiers and provide some rudimentary analysis of solved and open problems in this important domain. Our main contribution is a formal introduction of the notion of unverifiability, for which the paper could serve as a general citation in domains of theorem proving, as well as software and AI verification.","abs_url":"https://arxiv.org/abs/1609.00331","pdf_url":"https://arxiv.org/pdf/1609.00331v3","match":"abstract"},{"id":"1607.07730","version":1,"title":"A Model of Pathways to Artificial Superintelligence Catastrophe for Risk and Decision Analysis","authors":["Anthony M. Barrett","Seth D. Baum"],"published":"2016-07-25","updated":"2016-07-25","primary_category":"cs.AI","categories":["cs.AI"],"abstract":"An artificial superintelligence (ASI) is artificial intelligence that is significantly more intelligent than humans in all respects. While ASI does not currently exist, some scholars propose that it could be created sometime in the future, and furthermore that its creation could cause a severe global catastrophe, possibly even resulting in human extinction. Given the high stakes, it is important to analyze ASI risk and factor the risk into decisions related to ASI research and development. This paper presents a graphical model of major pathways to ASI catastrophe, focusing on ASI created via recursive self-improvement. The model uses the established risk and decision analysis modeling paradigms of fault trees and influence diagrams in order to depict combinations of events and conditions that could lead to AI catastrophe, as well as intervention options that could decrease risks. The events and conditions include select aspects of the ASI itself as well as the human process of ASI research, development, and management. Model structure is derived from published literature on ASI risk. The model offers a foundation for rigorous quantitative evaluation and decision making on the long-term risk of ASI catastrophe.","abs_url":"https://arxiv.org/abs/1607.07730","pdf_url":"https://arxiv.org/pdf/1607.07730v1","match":"both"},{"id":"1607.00913","version":1,"title":"Superintelligence cannot be contained: Lessons from Computability Theory","authors":["Manuel Alfonseca","Manuel Cebrian","Antonio Fernandez Anta","Lorenzo Coviello","Andres Abeliuk","Iyad Rahwan"],"published":"2016-07-04","updated":"2016-07-04","primary_category":"cs.CY","categories":["cs.CY","cs.AI"],"abstract":"Superintelligence is a hypothetical agent that possesses intelligence far surpassing that of the brightest and most gifted human minds. In light of recent advances in machine intelligence, a number of scientists, philosophers and technologists have revived the discussion about the potential catastrophic risks entailed by such an entity. In this article, we trace the origins and development of the neo-fear of superintelligence, and some of the major proposals for its containment. We argue that such containment is, in principle, impossible, due to fundamental limits inherent to computing itself. Assuming that a superintelligence will contain a program that includes all the programs that can be executed by a universal Turing machine on input potentially as complex as the state of the world, strict containment requires simulations of such a program, something theoretically (and practically) infeasible.","abs_url":"https://arxiv.org/abs/1607.00913","pdf_url":"https://arxiv.org/pdf/1607.00913v1","match":"both"},{"id":"1405.3378","version":1,"title":"The \"crisis of noosphere\" as a limiting factor to achieve the point of technological singularity","authors":["Rafael Lahoz-Beltra"],"published":"2014-05-14","updated":"2014-05-14","primary_category":"cs.CY","categories":["cs.CY"],"abstract":"One of the most significant developments in the history of human being is the invention of a way of keeping records of human knowledge, thoughts and ideas. In 1926, the work of several thinkers such as Edouard Le Roy, Vladimir Vernadsky and Teilhard de Chardin led to the concept of noosphere, thus the idea that human cognition and knowledge transforms the biosphere coming to be something like the planet's thinking layer. At present, is commonly accepted by some thinkers that the Internet is the medium that brings life to noosphere. According to Vinge and Kurzweil's technological singularity hypothesis, noosphere would be in the future the natural environment in which 'human-machine superintelligence ' emerges after to reach the point of technological singularity. In this paper we show by means of a numerical model the impossibility that our civilization reaches the point of technological singularity in the near future. We propose that this point may be reached when Internet data centers are based on \"computer machines\" to be more effective in terms of power consumption than current ones. We speculate about what we have called 'Nooscomputer' or N-computer a hypothetical machine which would consume far less power allowing our civilization to reach the point of technological singularity.","abs_url":"https://arxiv.org/abs/1405.3378","pdf_url":"https://arxiv.org/pdf/1405.3378v1","match":"abstract"}]}