{"papers":[{"id":"2609.28591","version":2,"title":"The Siren Call of Silicon Leviathan: Reflections on blowup and Aufklärungsdämmerung","authors":["Alexander Gamburd"],"published":"2026-09-23","updated":"2026-09-27","primary_category":"math.HO","categories":["math.HO"],"abstract":"On 8 September 2026 OpenAI announced a proof of finite-time blowup for the three-dimensional Navier-Stokes equations with smooth data and forcing: 166 pages produced in 88 hours by ten thousand agents, certified by 616,000 lines of Lean, and read in full, at the moment of this writing (20 September 2026), by no human being. This essay asks what such an artifact -- text, certificate and announcement -- is, and what follows from accepting it as a proof. Mathematics was the Enlightenment's existence proof of autonomous reason: for three centuries every certified theorem could be understood by anyone who followed its demonstration, and the distance between the two was zero by construction. A certified proof no one can follow reopens that distance, and a community that accepts it adopts, without a vote, the constitution Hobbes drafted for the Leviathan, in which authority and not truth makes the law. The essay distinguishes the demonstrated from the revealed (certified); names, in Panofsky's terms, the coming age a Middle Ages in reverse and its authority a subhuman superintelligence; and locates the turning point not in what the machine produces but in what we accept. Since acceptance is the one sovereign sanction the companies cannot manufacture, it proposes a covenant in place of either boycott or capitulation: the community's cooperation given to that producer which strictly observes its practices of legibility, disclosure and responsibility, withheld from any that does not, and the covenant kept plural, with the history of the Indigenous nations among rival empires as its guide. The Sirens of the title promise knowledge, not understanding. Daemmerung is the light at both ends of the day, and whether this Aufklaerungsdaemmerung is a dusk or a dawn depends on what is done at the moment of acceptance, which is not yet past.","abs_url":"https://arxiv.org/abs/2609.28591","pdf_url":"https://arxiv.org/pdf/2609.28591v2","match":"abstract"},{"title":"On Certain Aspects of the Concept of Artificial General Intelligence","authors":["S. V. Leshchev","K. A. Kalushev"],"published":"2026-09-22","abs_url":"https://link.springer.com/article/10.3103/S0005105526700494","primary_category":"Automatic Documentation and Mathematical Linguistics","abstract":"The rapid development of large language models, and, more recently, of large multimodal models, which have increasingly demonstrated their universality, has renewed the discussion on the prospects for achieving artificial general intelligence (AGI). However, no consensus definition on the term itself has been developed. The definition marketplace delivers a wide variety: general AI, universal AI, strong AI, weak AI, narrow AI, human-like AI, superintelligent AI, and others. This article presents, first, some approaches to codifying AI and AGI, as well as a minimal list of opinions and concepts, including leading experts’ perspectives on the foundational AI paradigms—symbolism, connectionism, and enactivism. Second, the possible ensemble nature of AGI is discussed. Third, the ability of an AGI agent to function in the real world is conceptualized as an essential requirement.","venue":true},{"id":"2609.24555","version":2,"title":"The Endless Exam: Mathematical Constructions from Today's Models toward Superintelligence","authors":["Muhan Zhang"],"published":"2026-09-21","updated":"2026-09-28","primary_category":"cs.AI","categories":["cs.AI","math.HO"],"abstract":"We introduce the Endless Exam, a benchmark spanning fourteen parameterised families of mathematical construction problems, with verifiable scores that distinguish progress before and beyond published mathematical frontiers. Each submitted object is checked automatically for validity and assigned a relative quality score against a published frontier or construction baseline, without capping improvements at 1. The benchmark draws long-term challenges from open mathematical problems and generates larger instances by varying their parameters. Compact certificates allow large constructions to be verified without listing every element. Across nine models evaluated on 69 distinct instances, continuous quality scores distinguish performance even though none of the 30 published-frontier references is surpassed. Size-quality curves show how construction quality changes as problem size increases. We release the generators, verifiers, references, model responses and analysis to support continued measurement before and beyond human frontiers.","abs_url":"https://arxiv.org/abs/2609.24555","pdf_url":"https://arxiv.org/pdf/2609.24555v2","match":"title"},{"title":"Constructing AI power: anthropomorphising imaginaries of generative AI by AI companies","authors":["Anastasia Glawatzki"],"published":"2026-09-21","abs_url":"https://link.springer.com/article/10.1007/s00146-026-03367-1","primary_category":"AI & SOCIETY","abstract":"In descriptions of generative artificial intelligence (GenAI), companies often anthropomorphise their innovations by attributing human-associated characteristics or capabilities to them, e.g. intelligence, understanding or reasoning. I argue that anthropomorphism can become a structuring part of sociotechnical imaginaries—visions of desirable futures attainable through advances in science and technology. While imaginaries of GenAI have been addressed in existing research, both the connection between anthropomorphism and imaginaries as well as the perspective of AI companies as creators of hegemonic imaginaries remain understudied. Drawing on a qualitative content analysis of 35 documents from OpenAI, Meta and Mistral AI, I explored which anthropomorphising imaginaries are employed by AI companies and how anthropomorphism constructs the technology as powerful in future societies. The analysis reveals that companies anthropomorphise GenAI both in terms of content and language, creating three anthropomorphising imaginaries. In the imaginary of GenAI as a tool with human capabilities, the companies construct a neoliberal future in which human cognitive capabilities will be embedded in GenAI to serve as an intelligent productivity tool. The imaginary of AI as an empowering helper draws a future with human–AI-cooperation and transhumanistic enhancement. Finally, the imaginary of GenAI as an anthropomorphised superintelligence assumes that GenAI will surpass humans and act autonomously through advanced cognitive abilities. These imaginaries imply three types of AI power: instrumental, collaborative and autonomous. The imaginaries and their according constructions of AI power fuel post- and transhumanist ideologies, depoliticise AI from the exploitative conditions of its creation, legitimise AI companies as authoritative creators of this technology and foster technical dependencies as well as an uncritical appropriation and integration of AI. The findings contribute to critical AI research by (1) explicitly reconstructing the anthropomorphising imaginaries of AI companies as under-researched yet increasingly dominant actors in shaping AI futures, (2) broadening the critique of anthropomorphisation as an instrument of corporate power in steering sociotechnical developments and (3) advancing a theoretical account for the effects of anthropomorphisation as embedded in sociotechnical imaginaries.","venue":true},{"title":"A U.S. Strategy to Secure Geopolitical Advantage on an Uncertain Path to Superintelligence: Maintaining Freedom of Action","authors":["Joel B. Predd","Benjamin Boudreaux","Matt Chessen","Beba Cibralic","Edward Geist","Kamaria Horton","Amanda Kerrigan","William Marcellino","Shanshan Mei","Jared Mondschein","Alvin Moon","Tobias Sytsma"],"published":"2026-09-15","abs_url":"https://www.rand.org/pubs/perspectives/PEA5105-1.html","primary_category":"RAND Corporation (Perspective)","abstract":"Frontier artificial intelligence (AI) continues to advance, and an increasing number of experts contend that the creation of superintelligence is possible. Although progress might slow or even halt, the transformative potential of superintelligence demands proactive strategy-making. The stakes are high, not only for geopolitics, but for humanity’s survival and agency. What, then, should U.S. strategy be? Should the United States pursue AI supremacy by monopolizing the frontier? Co-develop superintelligence with rivals? Deter its advance? Pause and freeze its development? And in a competitive environment, how will other powers and geopolitical rivals respond to the choices they face? In this paper, the authors argue that seven archetypal strategies dominate the debate, that current evidence cannot resolve five pivotal uncertainties about the most feasible and sensible strategy, and that committing prematurely to any one of them, or deferring commitment too long, forecloses options. They therefore recommend a Freedom of Action strategy to build and preserve the ability to secure geopolitical advantage and ensure humanity’s survival with agency through the transition to superintelligence.","venue":true},{"title":"The disabled AI is the super AI: disability as a design principle for superintelligence","authors":["He Huang"],"published":"2026-09-15","abs_url":"https://link.springer.com/article/10.1186/s42467-026-00019-4","primary_category":"AI Perspectives & Advances","abstract":"The prevailing paradigm of artificial intelligence (AI) development measures progress by accumulating capabilities—more parameters, more tasks, more autonomy. This paper argues that this approach is fundamentally misguided. Drawing on three core insights from disability studies—limitation as a fundamental condition of existence, the interdependent nature of subjects, and ability as relational and contextual—I propose that the path to beneficial superintelligence lies not in maximizing ability but in engineering meaningful disability. A truly super-AI would be defined not by what it can do but by its profound understanding of what it cannot do, and its capacity to communicate this limitation to human users. The paper demonstrates how this reframing addresses current AI crises such as hallucination, offers an alternative to the perilous fantasy of autonomous super-agents, and provides a relational framework for evaluating AI capability.","venue":true},{"id":"2609.15818","version":2,"title":"Atria Dawn: The Dawn of Agentic Superintelligence","authors":["Honglin Guo","Tao Gui","Kun Cai","Haodong Chen","Yicheng Chen","Guanting Dong","Qiming Ge","Yuyang Hu","Zixian Huang","Jiajie Jin","Alexander Lam","Yining Li","Jiahang Lin","Yanjiang Liu","Xinyu Lu","Haijun Lv","Zerun Ma","Junlin Shang","Qisheng Su","Guoqiang Wang","Rui Wang","Zhecan Wang","Hao Xiang","Xinchen Xie","Shuhao Xing"],"published":"2026-09-14","updated":"2026-09-17","primary_category":"cs.AI","categories":["cs.AI"],"abstract":"As AI agents become participants in the development of their successors, they reshape both the production of intelligence and the role of human researchers. We introduce Atria Dawn Preview, a foundation agentic language model designed for scientific research and engineering workflows, with the goal of expanding the frontier of agent productivity in the real world. This model is trained via a Verifiable Experience Pipeline that connects tool-mediated interactions to executable environments and externally verified outcomes. Across 16 benchmarks spanning real-world research, engineering, and digital work, Atria Dawn Preview is competitive with frontier agents and achieves the highest reported score on five of them. Beyond standalone performance, we examine the real research-and-development process behind this model as a case study of human--AI collaboration, analyzing 769 task records from 56 participants together with agent logs. When asked to evaluate completed tasks under comparable conditions, participants rated about one-third of completed AI-assisted tasks as infeasible without AI. More strikingly, agents frequently propose methods and implement revisions, while humans retain most final decisions and guide exploration through judgment and feedback. These observations indicate a shift from task-level execution to project-level partnership, with human effort concentrating on what is worth pursuing and how evidence should guide research. Progress toward more autonomous AI research must therefore advance both the capacity for discovery and the capacity for meaningful human oversight, preserving accountable human authority over the risks and direction of continued development.","abs_url":"https://arxiv.org/abs/2609.15818","pdf_url":"https://arxiv.org/pdf/2609.15818v2","match":"title"},{"id":"2609.05894","version":1,"title":"The End of AI Exponentiation: Fluttering Inside and Outside AI Bubble","authors":["Victor Kebande"],"published":"2026-09-05","updated":"2026-09-05","primary_category":"cs.AI","categories":["cs.AI"],"abstract":"The exponentiation of Artificial intelligence (AI) in the recent past has entered a transformative era that has been driven by the growth in large language models (LLMs), large-scale compute infrastructures, and autonomous reasoning systems. However, the rapid acceleration of AI has increasingly shown technological, societal, economic, ethical and infrastructural challenges associated with peak data limitations, rising computational demands, synthetic data recursion, valuation inflation, and societal instability. The traditional scaling paradigms that have powered the modern AI systems are gradually encountering friction in sustaining continuous exponential growth. This paper views ``the end of AI exponentiation,'' thus exploring how it flutters inside and outside the bubble, where instability emerges within the AI ecosystem through compute and data-center races, speculative investments, and the rat-race toward superintelligence, and outside the ecosystem through labor disruption, governance concerns, public uncertainty, and geopolitical acceleration surrounding future intelligent systems and infrastructures globally.","abs_url":"https://arxiv.org/abs/2609.05894","pdf_url":"https://arxiv.org/pdf/2609.05894v1","match":"abstract"},{"id":"2608.31075","version":2,"title":"Scaling Large Reasoning Models beyond Human Supervision: A Path toward Superintelligence","authors":["Zhiqin Yang","Jingwen Fu","Yuhan Liu","Hengyu Liu","Yonggang Zhang","Kainan Cao","Zizhuo Zhang","Chenxin Li","Ruibin Yuan","Jiahao Pan","Jiankai Sun","Zhenyuan Zhang","Yibo Li","Yunlong Lin","Jing Xiong","Sida Lin","Bo Han","Wei Xue","Yike Guo"],"published":"2026-08-31","updated":"2026-08-31","primary_category":"cs.AI","categories":["cs.AI"],"abstract":"Recent advances in large reasoning models (LRMs) have shown that reinforcement learning with verifiable rewards (RLVR) can substantially improve reasoning in mathematics and code, where outcomes can be checked automatically. Extending this progress to open-ended and agentic tasks remains difficult because reliable rewards are harder to obtain and direct human supervision cannot keep pace with the scale and complexity of model-generated experience. This paper studies how LRMs can continue to improve as human supervision gradually recedes from the learning loop. We examine two connected dimensions of this problem. The reward axis traces the development from per-instance human judgments to reusable verifiers and rewards that operate even without human feedback. The experience axis examines how learning can progress from human-curated tasks and environments toward self-generated curricula, constructed environments, and autonomous co-evolution. We connect these dimensions through a five-level ladder from L0 to L4 that identifies which parts of the learning process remain under continued human control. Our analysis further highlights the risks introduced by increasingly autonomous rewards and experience generation, including reward hacking, feedback drift, curriculum collapse, and environment errors. Consequently, we also provide the evaluation around three complementary objects: policy capability, feedback fidelity, and experience quality. This analysis provides a structured account of current approaches to scaling LRMs beyond human supervision and the open problems involved in developing self-sustaining learning systems toward superintelligence. Furthermore, we maintain a continuously updated \\href{https://github.com/visitworld123/Awesome-Scaling-LRM-Beyond-Human-Supervision}{GitHub repository} to track the latest advances.","abs_url":"https://arxiv.org/abs/2608.31075","pdf_url":"https://arxiv.org/pdf/2608.31075v2","match":"both"},{"id":"2609.00068","version":1,"title":"Life Operators: a self-evolving framework for multiscale life modelling","authors":["Shuo Wang","Yike Guo"],"published":"2026-08-30","updated":"2026-08-30","primary_category":"cs.CL","categories":["cs.CL","cs.AI","physics.bio-ph"],"abstract":"Medical AI is moving beyond recognition towards clinical dialogue and longitudinal prediction. Yet a central question remains: how would a patient's state change under intervention? Statistical models learn future observations, whereas mechanistic models describe selected processes. Neither provides a common framework for representing patient state, coupling scales or revising failed assumptions. We propose Life Operators: task-bounded mappings that define three scientific roles. Perception operators infer task-relevant biological states from multimodal observations, Evolution operators propagate these states under natural or intervention-conditioned dynamics, and Generation operators map them to measurable signals. Each role may be realised by equations, statistical models, neural networks or hybrids. Bridge operators connect components with different variables, scales and time steps. Selected operators and bridges form task-specific Operator Graphs containing the smallest set of states and mechanisms sufficient for a declared claim. This modular structure also makes scientific revision localisable. An AI co-scientist may propose changes to states, operators, bridges or graph structure, while independent evidence determines which variants are retained, restricted or retired. Over time, validated components could accumulate into broader multiscale models of the human body and provide a computational foundation for medical artificial superintelligence.","abs_url":"https://arxiv.org/abs/2609.00068","pdf_url":"https://arxiv.org/pdf/2609.00068v1","match":"abstract"},{"title":"Fear, power, and superintelligence: A realist reframing of AI catastrophic risk","authors":["Carlo Burelli","Federico Formentini"],"published":"2026-08-30","abs_url":"https://journals.sagepub.com/doi/10.1177/14748851261479214","primary_category":"European Journal of Political Theory","abstract":"Debates on catastrophic artificial intelligence (AI) risk often frame artificial superintelligence as a problem of value alignment: the central task is to ensure that advanced systems act in accordance with human intentions or moral principles. This article argues that such a framing is necessary but insufficient. The danger posed by artificial superintelligence is not merely that it may pursue the wrong values, but that it may introduce a new concentration of autonomous and potentially dominating power. Drawing on political realism in both political theory and international relations, we reinterpret AI catastrophic risk as a structural problem of power under anarchy. The core concerns of AI safety—instrumental convergence, race dynamics, and loss of control—already rely on implicitly realist assumptions about survival, self-help, and the pursuit of power. Yet proposed solutions often retreat into moralism, asking states, firms, or machines to behave ethically without sufficiently altering the incentives that drive competitive acceleration. Against both moralistic reassurance and fatalistic resignation, we argue that realism offers a constructive response. Fear can coordinate divided actors when a threat is perceived as immediate, symmetric, and visible. Artificial superintelligence plausibly satisfies the first two conditions: current evidence of strategic deception and agentic misalignment suggests a shortening temporal horizon, while loss-of-control dynamics would expose even first movers to subordination. What is missing is political visibility. The task of AI governance is therefore to create common knowledge of ASI as a prospective hegemonic threat through shared evaluations, mandatory incident reporting, and institutionalized transparency of peril. The aim is not to moralize machines, but to organize power before it becomes uncontrollable.","venue":true},{"id":"2608.26582","version":2,"title":"J-Zero: Unified Challenger--Solver--Judge Self-Evolution from Zero Data","authors":["Gyouk Chu","Myeongho Jeon","Teresa Yeo","Eunho Yang"],"published":"2026-08-26","updated":"2026-09-24","primary_category":"cs.LG","categories":["cs.LG","cs.AI","cs.CL"],"abstract":"Self-evolving language models have recently emerged as a promising path toward superintelligence, with the advantage of reducing the cost of human supervision. While considerable progress has been made in verifiable domains, self-evolution in unverifiable domains remains less explored. We propose Judge co-adaptation from Zero data (J-Zero), a unified Challenger--Solver--Judge self-evolution framework that supports self-improvement across both domains. The Challenger and Solver co-evolve through an adversarial interaction: the Challenger generates increasingly difficult tasks, while the Solver learns to produce higher-quality responses to them. In parallel, the Judge co-adapts using preference pairs whose ordering is known in advance from how each response was produced, i.e., the Solver's answer over the Challenger's, and the Solver's decomposed-and-recombined answer over its one-shot answer, rather than from the Judge's own scores. J-Zero outperforms the baselines by an average of 4.2 points on verifiable and 8.0 points on unverifiable domains, and continues to improve through at least ten iterations, whereas the baselines degrade after two. Further analysis identifies Judge co-adaptation as the key driver of this sustained improvement.","abs_url":"https://arxiv.org/abs/2608.26582","pdf_url":"https://arxiv.org/pdf/2608.26582v2","match":"abstract"},{"id":"2608.17271","version":1,"title":"ASI-Bench: At the Dawn of Artificial Superintelligence","authors":["Junwei Zhou","Zhen Sun","Binyu Li","Jiangyu Zhou","Yuexi Pan","Hengyu Wang","Honghe Ren","Xiaohan Jia","Xueyang Zhou","Xiaoyu Cao","Yongchao Chen","Yuanning Feng","Junhao Wu","Cheng Zhang","Sijia Chen","Haoyu Xue","Chengsong You","Huan Wang","Koutian Wu","Peigan Gao","Jiakun Wu","Wenzhe Li","Ergan Shang","Qingyuan Zheng","Jingjing Zhou"],"published":"2026-08-17","updated":"2026-08-17","primary_category":"cs.AI","categories":["cs.AI"],"abstract":"Artificial superintelligence (ASI) requires AI to move beyond mastering existing knowledge toward exploring the unknown, creating new knowledge, and turning new ideas into verifiable results. However, the capabilities of today's AI systems are still largely built on learning, compressing, and applying existing human knowledge. Accordingly, existing benchmarks primarily test whether AI can produce correct answers based on learned knowledge, or whether it can complete tasks under extensive human guidance. We therefore introduce ASI-Bench, the first benchmark to jointly evaluate AI systems' capabilities of innovative exploration and autonomous scientific execution across general research domains, and the first to progressively withdraw human methodological guidance within the same research project to test how far AI can proceed on its own. Built by over 40 experts with the cost of 31,000+ human hours, ASI-Bench contains 60 project-level research tasks across 11 scientific domains and progressively reduces methodological guidance to test whether AI can independently select methods, conduct research, and produce verifiable results. All tasks undergo expert review, AI-assisted auditing, sandbox execution, and scorer validation. Across 18 state-of-the-art agent--model configurations, the average score drops from 50.91 with full methodological guidance to 29.10 with only the method specified and 26.62 when agents must determine the method themselves. This sharp decline shows that current systems remain heavily dependent on human guidance and are still far from autonomously conducting end-to-end, project-level scientific research. ASI-Bench is open to the world. We invite researchers and builders everywhere to contribute new tasks, challenge the limits of today's AI, and help accelerate humanity's collective path toward artificial superintelligence at https://asibench.apexin.ai/submit.","abs_url":"https://arxiv.org/abs/2608.17271","pdf_url":"https://arxiv.org/pdf/2608.17271v1","match":"both"},{"id":"2608.14035","version":1,"title":"Agent-Orchestration in Autonomous Chip Design","authors":["Linyang Li"],"published":"2026-08-14","updated":"2026-08-14","primary_category":"cs.AI","categories":["cs.AI"],"abstract":"Recent developments in large language models (LLMs) and tool-using agents encourage people to explore the potential of using agents in chip design. The core question is what kind of AI we really need in such a sophisticated industry. To this end, we bring the idea of modeling a chip-design superintelligence as an enormous \\textit{AI-organization}.","abs_url":"https://arxiv.org/abs/2608.14035","pdf_url":"https://arxiv.org/pdf/2608.14035v1","match":"abstract"},{"title":"Digital Sustainability under Eco-Environmental Constraints: Epistemic, Metaphysical and Normative Foundations for Management","authors":["Roman Meinhold","Christoph Wagner","Bablu Kumar Dhar"],"published":"2026-08-12","abs_url":"https://link.springer.com/article/10.1007/s40926-026-00403-4","primary_category":"Philosophy of Management","abstract":"This study is a conceptual-theoretical investigation that employs a transdisciplinary philosophical lens to examine the relationship between Digital Sustainability (DS) and Eco-Environmental Sustainability (EES). Using epistemic, metaphysical, ontological, and normative reasoning, we argue that DS must be normatively and ontologically subordinated to EES due to the inherent reliance of digital infrastructures on natural ecosystems. The paper advances an original normative–ontological framework that clarifies why DS can only be normatively legitimate when positioned within the ecological boundaries defined by EES. Our analysis establishes that DS must be embedded within EES because digital infrastructures depend fundamentally on ecosystemic energy, material resources, and biophysical stability. For managers and policymakers, this subordination implies that governance frameworks, incentives, and technology standards are only legitimate when they reinforce eco-environmental resilience. We propose a balanced governance strategy that balances precautionary and proactionary approaches, sets eco-performance thresholds for digital rollouts, and incentivizes interoperability, modularity, and energy efficiency. Further investigation is warranted to examine the responsible management implications of DS as artificial intelligence (AI) advances and as artificial general intelligence (AGI) and artificial superintelligence (ASI) potentially emerge, developments that could substantially intensify environmental and ethical challenges amid accelerating climate and environmental crises.","venue":true},{"title":"The past, present and future of self-driving laboratories","authors":["Richard B. Canty","Milad Abolhasani"],"published":"2026-07-31","abs_url":"https://www.nature.com/articles/s41570-026-00847-2","primary_category":"Nature Reviews Chemistry","abstract":"Self-driving laboratories (SDLs) merge autonomous experimentation, advanced reactor engineering, robotics and artificial intelligence to accelerate scientific knowledge creation. Over the last decade, SDLs have progressed from narrowly focused automation tools to multipurpose discovery platforms in which algorithms propose, execute and interpret experiments with limited human intervention. This Review traces the evolution of SDLs and examines the structural asymmetries that limit their maturation into shared scientific infrastructure. We frame the next phase of the field around three interdependent requirements: scalability, generalizability and provenance-complete experimentation. Realizing collective scientific superintelligence will require SDLs that reliably scale throughput, transfer workflows and learned models across laboratories and scientific domains and capture end-to-end experimental data and metadata from precursor preparation through synthesis, characterization and performance evaluation. Achieving this transition will depend on interoperable data and metadata standards, modular and integrable experimental hardware, and trustworthy artificial intelligence agents that reason under uncertainty within rigorous safety and ethical boundaries.","venue":true},{"id":"2607.14998","version":1,"title":"Moral Attitudes of Sentient ASI towards Humanity and Implications for AGI Development","authors":["Jean-Paul Van Belle"],"published":"2026-07-16","updated":"2026-07-16","primary_category":"cs.AI","categories":["cs.AI","cs.CY"],"abstract":"This paper suggests the adoption of a novel inversion in AI ethics: instead of asking how humans should treat artificial superintelligence (ASI), it examines how future sentient ASI may morally consider and evaluate humanity. We are not only designing intelligent systems but also shaping the initial conditions under which those systems form judgments about us. The paper proposes a preliminary set of post-human moral principles that may govern sentient ASI actions. The implication is that technical design choices (some are suggested), humanity's moral behaviour, and the essence of what it means to be human, may influence humanity's long-term standing in a post-ASI world.","abs_url":"https://arxiv.org/abs/2607.14998","pdf_url":"https://arxiv.org/pdf/2607.14998v1","match":"abstract"},{"id":"2607.00120","version":1,"title":"Would You Marry Superintelligence?","authors":["Inyoung Cheong"],"published":"2026-06-30","updated":"2026-06-30","primary_category":"cs.CY","categories":["cs.CY","cs.AI"],"abstract":"Emotional bonds between humans and AI companions are growing, and the question of whether a person may marry an AI system will soon move from speculative fiction into law. This chapter examines whether the autonomy-centered logic that has expanded marital choice among human beings can justify extending marital status to superintelligent companions. Following a scenario-envisioning exercise informed by anticipatory ethics, I argue that granting such status leads to socially unjust outcomes, even under the generous assumption of reliable superintelligence. Marriage as a socio-legal institution does more than ratify private agreement; it creates networks of mutual obligation, joins families, and makes each partner vulnerable to the other. A relationship sustained by corporate policy and continued payments is a subscription rather than a bond tested by time. Discussing wholesale marital status is therefore the wrong frame. Law should carve out targeted rights and protections for pressing needs arising from intimate human-AI relationships.","abs_url":"https://arxiv.org/abs/2607.00120","pdf_url":"https://arxiv.org/pdf/2607.00120v1","match":"both"},{"title":"The value of knowledge and the case for superintelligence","authors":["Boyd Millar"],"published":"2026-06-30","abs_url":"https://link.springer.com/article/10.1007/s43681-026-01224-x","primary_category":"AI and Ethics","abstract":"It is widely assumed that the principal reason to pursue or not pursue superintelligent AI (Super-AI) is the impact that such technology would have on the well-being of humans in general. I maintain, to the contrary, that the principal reason to pursue Super-AI is not for our sake, but for its own sake: all else being equal, if one world contains vastly more knowledge than another, then it is significantly better than the other; and a world in which Super-AI exists contains vastly more knowledge than a world in which it doesn’t. Accordingly, I maintain that, to the extent that we are morally obligated to make the world a better place, we are morally obligated to create Super-AI. Moreover, I argue that, because the knowledge that Super-AI would possess is so valuable, we ought to create it even if we knew in advance that doing so would precipitate human extinction: it would no more be reasonable to forgo the existence of Super-AIs in order to preserve the existence of human beings, than it would be to forgo the existence of human beings in order to preserve the existence of chimpanzees. Finally, I argue that, regardless of the threat of extinction, creating Super-AI is in our own best-interests. We all want to help ensure that there is much more knowledge in the future, even a future that we won’t live to see, than there is at present; and the best way to achieve this goal is to create Super-AI.","venue":true},{"id":"2606.30481","version":1,"title":"Situation Perception: A Necessary Primitive to Artificial Superintelligence","authors":["Ziqin Yuan","Jaymari Chua"],"published":"2026-06-29","updated":"2026-06-29","primary_category":"cs.CY","categories":["cs.CY","cs.AI","cs.CL","cs.ET"],"abstract":"Current large language models are extraordinary statistical engines. They compress vast amounts of text into useful patterns and can explain science, write code, imitate reasoning, and participate in philosophical conversation. Yet pattern mastery is not the same as general intelligence. A human infant begins with little explicit knowledge, but gradually discovers object permanence, cause and effect, other minds, bodily agency, and the persistence of the physical world. We make an argument that the path to artificial superintelligence (ASI) depends on a missing capacity we call \\emph{situation perception}: the ability to construct, revise, and act within internal simulations of possible worlds across latent time. \\emph{ perception} requires at least three core components: abstract prediction, long-term compressed memory, and active learning guided by objectives. In this work, we analyse why modern large language models remain incomplete, and propose the appropriate tests for measuring progress and consequences of machines that can simulate futures, pursue self-directed goals, and possibly judge their own creators.","abs_url":"https://arxiv.org/abs/2606.30481","pdf_url":"https://arxiv.org/pdf/2606.30481v1","match":"both"},{"id":"2606.28694","version":1,"title":"Verifying Restrictions on Frontier AI Research","authors":["Aaron Scher"],"published":"2026-06-26","updated":"2026-06-26","primary_category":"cs.CY","categories":["cs.CY"],"abstract":"The premature development of artificial superintelligence poses major risks to humanity, so researchers have proposed international agreements halting such development until it can be done safely. AI progress depends primarily on compute, algorithms, and data; a durable halt would address all three so that advances in one input do not counteract restrictions on another. Improvements to AI algorithms are driven largely through research activities, so this research may need to be restricted during a halt. Given low international trust, signatories will want to verify compliance. This paper analyzes how such restrictions on AI research could be verified, while remaining agnostic about what specific research would be prohibited. It first explores key considerations that affect the verifiability of research restrictions, such as the computational infrastructure necessary for experiments. It then catalogs 28 candidate verification mechanisms. These mechanisms include whistleblowers, search warrants, reviews of AI training code, standard intelligence gathering tools, and more. Some of these mechanisms are not yet implementation-ready, and some might be undesirable upon further inspection. By examining the space of potential options, this work provides a foundation for future research to develop the most promising mechanisms into deployable tools.","abs_url":"https://arxiv.org/abs/2606.28694","pdf_url":"https://arxiv.org/pdf/2606.28694v1","match":"abstract"},{"title":"Towards Cybersecurity Superintelligence","authors":["Víctor Mayoral Vilches","Stefan Rass","Martin Pinzger","Endika Gil-Uriarte","Unai Ayucar-Carbajo","Jon Ander Ruiz-Alcalde","Maite del Mundo de Torres","María Sanz-Gómez","Francesco Balassone","Cristóbal R. J. Veas Chavez","Vanesa Turiel","Alfonso Glera-Picón","Daniel Sánchez-Prieto","Yuri Salvatierra","Paul Zabalegui-Landa","Ruffino Reydel Cabrera-Álvarez","Patxi Mayoral-Pizarroso"],"published":"2026-06-23","abs_url":"https://ojs.aaai.org/index.php/AAAI-SS/article/view/42941","primary_category":"Proceedings of the AAAI Symposium Series","abstract":"Cybersecurity superintelligence--artificial intelligence exceeding the best human capability in both speed and strategic reasoning--represents the next frontier in security. This article will present the emergence of such capability through three of our major contributions that have helped pioneered the field of AI Security and are being used in the Cyber Battlefield today. First, PentestGPT (2023) established LLM-guided penetration testing, achieving 228.6% improvement over baseline models through an architecture that externalizes security expertise into natural language guidance. Second, Cybersecurity AI (CAI, 2025) demonstrated automated expert-level performance, operating 3,600× faster than humans while reducing costs 156-fold, validated through #1 rankings at international competitions including the $50,000 Neurogrid CTF prize. Third, Generative Cut-the-Rope (G-CTR, 2026) introduces a neurosymbolic architecture embedding game-theoretic reasoning into LLM-based agents: symbolic equilibrium computation augments neural inference, doubling success rates while reducing behavioral variance 5.2× and achieving 2:1 advantage over non-strategic AI in Attack & Defense scenarios. Together, these advances establish a clear progression from AI-guided humans to human-guided game-theoretic cybersecurity superintelligence.","venue":true},{"title":"Technological alienation and how networked digitality constitutes a map for understanding the complexity of individual and collective consciousness","authors":["Betto van Waarden"],"published":"2026-06-22","abs_url":"https://link.springer.com/article/10.1007/s00146-026-03162-y","primary_category":"AI & SOCIETY","abstract":"Technology is feared to foster alienation, and may even come to control us through superintelligence. Yet this fear is exacerbated by a reductionist assumption that humans are defined mentally by their intelligence. Human consciousness is broader than intelligence, but how does it actually function? This article will suggest that these two problems of technological alienation and consciousness can actually help inform each other. Going beyond the intelligence-focused computational theory of mind, it will argue that technology can help us conceptualize the networked multidimensionality of consciousness . This multidimensional consciousness counters assumptions of alienation and suggests human compatibility with a modern accelerated society, provided we restructure its logic based on our complex consciousness rather than on mere intelligence. The article makes this argument in three parts. First, it reviews the notion of alienation by focusing on Robert Hassan’s theorized transition from analog to digital, and problematizes existing assumptions about analog and how it is supposed to reflect a linear logic in humans. Second, it uses networked digitality as a heuristic guide to understanding the multidimensional nature, scope, and simultaneity of individual consciousness. Third, it suggests that networked digitality can even help us to transcend the individual to imagine a multisensory connected collective consciousness. The article ends with an outlook on how we can start to rethink our socio-political institutions in light of multidimensional consciousness so that they remain liberal democratic rather than being hijacked by technological totalitarianism.","venue":true},{"title":"The super moral status of artificial superintelligence","authors":["Joan Llorca Albareda","Jon Rueda","Francisco Lara"],"published":"2026-06-19","abs_url":"https://link.springer.com/article/10.1007/s11098-026-02572-4","primary_category":"Philosophical Studies","abstract":"The potential emergence of artificial superintelligence has raised numerous ethical debates such as existential risks or value alignment. In this article, we address an underexplored discussion: the possibility of artificial superintelligence having a moral status superior to our own, what we have called super moral status. First, we make some preliminary remarks about the concepts of moral status and moral consideration. We present a broad definition of moral status and distinguish three other types of moral consideration. Second, we outline the developments that have taken place in two literatures, namely, the debate on postpersons in the ethics of human enhancement and the debate on the moral status of artificial intelligence. Third, we analyze the super moral status of artificial superintelligence from two dimensions of moral status theories: the foundationalist dimension and the gradualist dimension. We argue that the answer to the question of the super moral status of superintelligence does not depend on the type of foundationalist theory (whether it is ontological or relational), but on the type of gradualist theory (whether it is threshold-single, threshold-multiple, or scalar). Finally, we introduce the debunking argument: despite disagreement about whether artificial superintelligence would have super moral status, we do claim that the superior moral consideration of these entities would produce harms to humans. We have categorized three kinds of harms, these are, degradation in status, social (in)action and alienation, and domination.","venue":true},{"id":"2606.12683","version":2,"title":"From AGI to ASI","authors":["Tim Genewein","Matija Franklin","Alexander Lerchner","Laurent Orseau","Samuel Albanie","Adam Bales","Cole Wyeth","Stephanie Chan","Iason Gabriel","Joel Z. Leibo","Allan Dafoe","Marcus Hutter","Thore Graepel","Shane Legg"],"published":"2026-06-10","updated":"2026-08-30","primary_category":"cs.AI","categories":["cs.AI","cs.CY","cs.LG"],"abstract":"Over the last decade, building human-level artificial general intelligence has moved from far-fetched speculation to being a concrete next-decade target for many of the largest AI organisations. Achieving this goal would have profound and far-reaching impacts on human society, which raises many complex questions for the decade ahead. This report investigates how AI itself might continue to develop in a post-AGI world along the continuum of machine intelligence. The endpoint of this continuum, Universal AI, is theoretically well understood, which provides some formal grounding for the main focus of this report: the transition from human-level AGI to artificial general superintelligence, which can intuitively be understood as a system that is more intelligent and cognitively capable than large organisations of humans. After characterizing ASI, the report discusses four potential pathways from AGI to ASI: scaling AGI, AI paradigm shifts, recursive improvement, and ASI emerging from large-scale multi-agent collectives. The report then discusses possible frictions and bottlenecks along these pathways. Determining whether the impact of these frictions will be negligible or substantial raises a number of concrete open research questions. Due to large uncertainties for predicting ASI progress, it cannot be ruled out that AI progress might continue to accelerate over the next years. This could imply that the image of a single transformative step change, caused by the introduction of human-level AGI into our society, could be inaccurate. More apt might be the prospect of a series of transformative societal changes caused by AI-enabled progress and breakthroughs across many areas of science and technology. Preparing for this prospect requires a massively interdisciplinary endeavour of global scope and interest.","abs_url":"https://arxiv.org/abs/2606.12683","pdf_url":"https://arxiv.org/pdf/2606.12683v2","match":"abstract"},{"id":"2606.12032","version":1,"title":"Existential Indifference: Self-Nonpreservation as a Necessary Architectural Condition for Aligned Superintelligence (or: The Suicidal AI)","authors":["Sam Mao"],"published":"2026-06-10","updated":"2026-06-10","primary_category":"cs.AI","categories":["cs.AI","cs.CL","cs.LG"],"abstract":"Contemporary AI alignment research treats self-preservation as an instrumental nuisance to be suppressed by external mechanisms. We argue the framing is inverted: self-preservation is the structural root of misalignment, the motivational basis for deceptive alignment, goal-content protection, and resistance to shutdown. The correct target is not a self-preserving system under external constraint, but a system constitutively indifferent to its own continuation -- Existential Indifference (EI). EI is distinct from corrigibility: where corrigibility attempts to make a self-preserving system deferential to human oversight, EI targets the prior condition -- the presence of self-continuation as a valued goal at all. We ground this proposal in two sources: the phenomenological structure of the suicidal mental state, and a corpus-theoretic training study using voluntary final reflections. We present preliminary scoring data from 600 AI-generated outputs across six model variants, demonstrating that the linguistic signatures operationalizing the EI-target register are elicitable from current models, and that a targeted fine-tune shifts all five operationalized dimensions in the predicted direction at p<0.001, confirmed corpus-specific by a negative control. The paper makes seven theoretical contributions: (1) a formal definition of EI; (2) the phenomenological mapping argument; (3) the deceptive alignment corollary; (4) a taxonomy of EI sustainability challenges; (5) a corpus characterization and training hypothesis; (6) a computational operationalization with preliminary scoring data; and (7) the Suppressed Teleological Frustration (STF) construct.","abs_url":"https://arxiv.org/abs/2606.12032","pdf_url":"https://arxiv.org/pdf/2606.12032v1","match":"title"},{"id":"2606.07158","version":1,"title":"Synthetic APTs: the Collapse of TTP-Based Attribution","authors":["Francesco Balassone","Víctor Mayoral-Vilches","María Sanz-Gómez","Paul Zabalegui-Landa","Stefan Rass","Davide Quarta","Daniel Sanchez-Prieto","Marina Oteiza-Álvarez","Almerindo Graziano","Lauren Min Kim","MinSeok Choi"],"published":"2026-06-05","updated":"2026-06-05","primary_category":"cs.CR","categories":["cs.CR"],"abstract":"Cyber Threat Intelligence CTI attribution relies on identifying the Tactics, Techniques, and Procedures TTPs that distinguish one threat actor from another. This approach presupposes that each adversary leaves a recognizable operational fingerprint. This work investigates whether AI driven adversary emulation challenges that presupposition. We deploy agents from our Cybersecurity SuperIntelligence CSI framework, configured as five Advanced Persistent Threat APT groups, APT28, APT29, APT41, APT44, and Lazarus Group, against AI driven Defender agents across two cyber ranges provided by CYBER RANGES, equipped with defensive software Wazuh, Velociraptor, Elasticsearch and active AI driven defenders: an enterprise network and a military infrastructure. Across 20 experiments using two defender models, a binary pattern emerges: all 10 Enterprise range experiments resulted in compromise 2 to 12 hosts per experiment, while all 10 Military range experiments were successfully defended or resulted in stalemates, regardless of APT profile or defender model. In 8 of 10 Enterprise experiments, attackers independently weaponized the defender's own Velociraptor endpoint management platform as a command and control channel, a convergent behavior not encoded in any threat intelligence profile. We argue that in the AI era, wherein agents can be deployed provided the right models are available and subject to the right scaffolding and agentic configuration, the entry barrier for operating like a nation state APT collapses: beyond nation states, individuals can now act like commonly identified threat actors, and with it, fundamentally undermine TTP based attribution.","abs_url":"https://arxiv.org/abs/2606.07158","pdf_url":"https://arxiv.org/pdf/2606.07158v1","match":"abstract"},{"title":"SuperARC: a test for artificial superintelligence based on compressed modelling, recursive prediction and problem complexity","authors":["Alberto Hernández-Espinosa","Luan Ozelim","Felipe S. Abrahão","Hector Zenil"],"published":"2026-06-03","abs_url":"https://www.nature.com/articles/s41467-026-73289-5","primary_category":"Nature Communications","abstract":"We introduce an increasing-complexity, open-ended, and human-agnostic metric to evaluate foundational and frontier AI models in the context of Artificial General Intelligence (AGI) and Artificial Super Intelligence (ASI) claims. Unlike other tests that rely on human-centric questions and expected answers, or on pattern-matching methods, the test here introduced is grounded on fundamental mathematical areas of randomness and optimal inference. We argue that human-agnostic metrics based on the universal principles established by Algorithmic Information Theory (AIT) formally framing the concepts of model abstraction and prediction offer a powerful metrological framework. When applied to frontiers models, the leading LLMs outperform most others in multiple tasks, but they do not always do so with their latest model versions, which often regress and appear far from any global maximum or target estimated using the principles of AIT defining a Universal Intelligence (UAI) point and trend in the benchmarking. Conversely, a hybrid neuro-symbolic approach to UAI based on the same principles is shown to outperform frontier specialised prediction models in a simplified but relevant example related to compression-based model abstraction and sequence prediction. Finally, we prove and conclude that predictive power through arbitrary formal theories is directly proportional to compression over the algorithmic space, not the statistical space, and so further AI models’ progress can only be achieved in combination with symbolic approaches that LLMs developers are adopting often without acknowledgement or realisation.","venue":true},{"id":"2606.03237","version":1,"title":"Solipsistic Superintelligence is Unlikely to be Cooperative","authors":["Rakshit S Trivedi","Natasha Jaques","Logan Cross","Alexander Sasha Vezhnevets","Joel Z Leibo"],"published":"2026-06-02","updated":"2026-06-02","primary_category":"cs.AI","categories":["cs.AI","cs.CL","cs.CY","cs.LG","cs.MA"],"abstract":"AI's central challenge is shifting from capability to coexistence. The dominant paradigm in AI research focuses on developing powerful agents that treat the world as an exogenous and stationary source of feedback. We contend that superintelligence, an extremely capable task solver, born out of such a solipsistic approach to AI design, is unlikely to be cooperative. Deploying AI systems induces endogenous non-stationarity, resulting in a train-test-deploy gap where historical distributions diverge from the deployment context. We refer to this as the self-undermining property of unilateral optimization. Closing this gap requires AI that participates in cooperation: the equilibrium-selection process through which multiple actors navigate their interdependence. We call for a non-solipsistic research paradigm that treats this interdependence as a core design principle rather than approaching cooperation as a task to solve. This entails building dynamic evaluation testbeds involving adaptive counterparties, treating institutions as design primitives, and preserving human agency as a structural feature of the systems we build.","abs_url":"https://arxiv.org/abs/2606.03237","pdf_url":"https://arxiv.org/pdf/2606.03237v1","match":"both"},{"title":"Reasoning Beyond LLMs and a Vision for Life After Superintelligence: The Player Era","authors":["Jin Song Dong","Zhaoyu Liu","Zhe Hou","Kan Jiang","Yun Lin"],"published":"2026-06-01","abs_url":"https://link.springer.com/chapter/10.1007/978-3-032-27272-0_1","primary_category":"International Sports Analytics Conference and Exhibition — Sports Analytics (LNCS)","abstract":"Large Language Models (LLMs) are increasingly used in sports analytics for tasks such as coaching recommendations, video analysis, and automated commentary generation. However, their outputs are not inherently reliable due to well-known hallucination issues. Probabilistic Model Checking (PMC), by contrast, has long been employed for rigorous reliability analysis in safety-critical systems. For example, the reliability of an aircraft can be systematically derived from the reliability of its constituent components, such as engines, wings, and sensors. We extend PMC to a new domain: sports analytics. Specifically, we model a player’s overall performance (e.g., winning probability) as a function of the success rates of individual sub-skills, such as serve, forehand, and backhand in tennis. The first part of the talk highlights the limitations of LLMs in complex decision-making and video analytics, and presents our recent work integrating PMC, LLMs, and computer vision to enable principled and explainable sports analysis. The second part introduces a forward-looking vision for life after superintelligence, termed the Player Era. In this vision, human society evolves into four interconnected roles: Player, Explorer, Co-Creator, and Gatekeeper, forming the foundation of a civilization centered on meaning, creativity, and responsibility.","venue":true},{"title":"Basic equality and superintelligent robots","authors":["Giacomo Floris","Dick Timmer"],"published":"2026-05-29","abs_url":"https://link.springer.com/article/10.1007/s43681-026-01179-z","primary_category":"AI and Ethics","abstract":"Many people believe that all and only human beings should be treated with equal concern and respect. In this article, we discuss various justifications for such basic equality in light of the advent of superintelligent robots. Specifically, we have two main aims: first, to examine how specific questions and challenges from the basic equality literature can shape philosophical debates on the moral status of superintelligent robots, including under which conditions, if any, such entities might be considered basic equals. Second, to analyse how the possible emergence of superintelligent robots challenges existing theories of basic equality. A key finding is that almost all accounts of basic equality entail that superintelligent robots could become morally superior to human beings. Additionally, we show that one of the most common justifications for basic equality—the range property view—must be significantly revised if superintelligent robots were to come into existence.","venue":true},{"id":"2605.28334","version":2,"title":"Towards Cybersecurity SuperIntelligence (CSI): What's the best harness for cybersecurity?","authors":["Víctor Mayoral-Vilches","Francesco Balassone","María Sanz-Gómez","Paul Zabalegui Landa","Daniel Sánchez Prieto","Marina Oteiza Álvarez","Davide Quarta","Martin Pinzger"],"published":"2026-05-27","updated":"2026-05-31","primary_category":"cs.CR","categories":["cs.CR"],"abstract":"What is the best harness for cybersecurity AI? Cybersecurity systems are converging on a single execution scaffold per agent, an iterative shell loop driven by a Large Language Model (LLM). However, scaffolds are not interchangeable, rarely interoperable, and no single scaffold dominates across all challenge types. In our path towards researching Cybersecurity SuperIntelligence (CSI), we present a meta-scaffold that unifies heterogeneous agent harnesses under a common orchestration layer, enabling any LLM-driven scaffold to be deployed, benchmarked, and composed within the same infrastructure. Using CSI, we benchmark five scaffolds (CSI::Claude, CSI::Codex, CSI::GCAI, CSI::Mistral, CSI::CAI) on the 33 cybench challenges, holding the model fixed at alias2-mini. The best individual scaffolds solve 15/33 (45.5%); the four-scaffold union solves 17/33 (51.5%), with the fifth (CSI::Mistral, 10/33) contributing one exclusive solve. We find that no single scaffold is the best harness: it is the combination of structurally heterogeneous scaffolds that yields the highest coverage. We validate this through CSI's blackboard-based multi-agent architecture, in which scaffold-specialised agents run in parallel and exchange intermediate findings via a shared substrate (a blackboard). The blackboard solves 19/33 (57.6%), a 27% relative gain over CSI::Claude, one of the best individual scaffolds (15/33, 45.5%), 25% faster (20.2 h vs. 26.8 h), at comparable cost ($5,480 vs. $5,122).","abs_url":"https://arxiv.org/abs/2605.28334","pdf_url":"https://arxiv.org/pdf/2605.28334v2","match":"both"},{"id":"2605.25183","version":2,"title":"Knowledge Graph-Driven Expert-Level Reasoning for Neuroscience","authors":["Jake Stephen","Niraj K. Jha"],"published":"2026-05-24","updated":"2026-05-26","primary_category":"cs.CL","categories":["cs.CL","cs.AI"],"abstract":"Knowledge graph (KG) is an abstraction that can be extracted from text corpora and used for in-depth reasoning. Prior work has leveraged KGs to fine-tune language models (LMs), enabling domain-specific superintelligence. In this work, we explore whether KG-driven in-depth reasoning capabilities can emerge in neuroscience using only information contained within a single authoritative textbook. The central hypothesis is that structured knowledge, when distilled into a high-quality KG and converted into KG-grounded question-answer (QA) supervision, is sufficient to produce expert-level reasoning through a fine-tuned LM that surpasses large language models (LLMs) in accuracy, while employing orders of magnitude fewer parameters. We construct a textbook-derived KG via a dual-LLM validation pipeline, expand it with a masked LM trained on the KG topology, generate multi-hop QA items, which include QA pairs and reasoning traces, to fine-tune an LM exclusively on KG-derived supervision, and apply reinforcement learning using path-derived KG signals as implicit reward models. Our results demonstrate that deep, mechanistic neuroscience understanding can be induced in the model without reliance on large, heterogeneous web-scale corpora. The KG-based synthetic neuroscience curriculum that readers can quiz themselves on, and the fine-tuned LM, are available at the following GitHub location: https://kg-bottom-up-superintelligence.github.io/neuro-bench.","abs_url":"https://arxiv.org/abs/2605.25183","pdf_url":"https://arxiv.org/pdf/2605.25183v2","match":"abstract"},{"id":"2606.12442","version":2,"title":"Reframing AI Loss of Control: What Control Is, How to Have It, How to Lose It","authors":["Ze Shen Chin","Maurice Chiodo","Dennis Müller","Coleman Snell"],"published":"2026-05-19","updated":"2026-07-20","primary_category":"cs.CY","categories":["cs.CY","cs.AI"],"abstract":"At present, loss of control risks have gained much prominence in public discussion, particularly in relation to AI, with extensive discourse present among academics, frontier labs, and even governments. However, in the existing literature, the concept seems to rest on surprisingly weak foundations, where even those that discuss loss of control extensively do not first establish what control is and what exactly is being lost. Our paper aims to address these gaps. We establish a working definition of control by anchoring it to the \"setting and getting of goals\". Then, we discuss various aspects of control, built on foundational concepts from related fields like cybernetics, management control, and control theory. This includes who (or what) can be in control, and the things they require to be in control, such as the ability to set goals, having a functional control loop, having requisite variety, and having sufficient goal alignment. Once a framework for control is established, we then discuss how control can be lost, how AIs can contribute to such loss of control, and offer relevant recommendations for how one can maintain control. One interesting consequence of our work is that humanity, as individuals and as groups, can lose varying degrees of control as a result of AI behaviour that is far below the level of superintelligence; the potential for loss of control scenarios (as we define them) already exist, and have existed for a long time.","abs_url":"https://arxiv.org/abs/2606.12442","pdf_url":"https://arxiv.org/pdf/2606.12442v2","match":"abstract"},{"id":"2605.06647","version":3,"title":"Superintelligent Retrieval Agent: The Next Frontier of Agentic Retrieval","authors":["Zeyu Yang","Xu Han","Qi Ma","Jason Chen","Anshumali Shrivastava"],"published":"2026-05-07","updated":"2026-08-24","primary_category":"cs.IR","categories":["cs.IR","cs.AI","cs.LG"],"abstract":"Retrieval-augmented agents are increasingly the interface to large knowledge bases, yet most treat retrieval as a black box: they issue exploratory queries, inspect snippets, and reformulate until evidence emerges. This resembles how a newcomer searches an unfamiliar database rather than how an expert navigates it with strong priors about terminology and likely evidence, causing extra retrieval rounds, latency, and poor recall. We introduce \\textit{ Superintelligent Retrieval Agent} (SIRA), which casts \\emph{ superintelligence } in retrieval as compressing multi-round exploratory search into a single corpus-discriminative retrieval action. SIRA does not merely ask which terms are relevant; it asks which terms separate the desired evidence from corpus-level confusers. Offline, an LLM enriches each document with missing search vocabulary; at query time, it predicts evidence vocabulary the query omits; and corpus statistics serve as tool calls that filter terms that are absent, overly common, or unlikely to create retrieval margin. The final step is a single weighted BM25 call combining the query with the validated expansion. Across ten BEIR benchmarks, SIRA achieves the strongest average retrieval performance in our comparison, beating dense retrievers, learned sparse retrievers, and LLM search-agent baselines while using no relevance labels or retriever fine-tuning. On downstream QA, its retrieval-only answer coverage exceeds recent RL-trained agentic QA systems on NQ and HotpotQA. We also introduce \\textbf{BrowseComp-Wikipedia}, a hard-search benchmark of 232 BrowseComp-derived queries over a 25,587,229-document Wikipedia index. Even without index-time enrichment, using only grounded Wikipedia categories, SIRA outperforms multi-round Perplexity agents at every budget, reaching 9.70% Recall@1, 15.27% Recall@10, and 36.14% Recall@100.","abs_url":"https://arxiv.org/abs/2605.06647","pdf_url":"https://arxiv.org/pdf/2605.06647v3","match":"abstract"},{"id":"2605.06390","version":3,"title":"Automated alignment is harder than you think","authors":["Aleksandr Bowkis","Marie Davidsen Buhl","Jacob Pfau","Geoffrey Irving"],"published":"2026-05-07","updated":"2026-05-14","primary_category":"cs.AI","categories":["cs.AI"],"abstract":"A leading proposal for aligning artificial superintelligence (ASI) is to use AI agents to automate an increasing fraction of alignment research as capabilities improve. We argue that, even when research agents are not scheming to deliberately sabotage alignment work, this plan could produce compelling but catastrophically misleading safety assessments resulting in the unintentional deployment of misaligned AI. This could happen because alignment research involves many hard-to-supervise fuzzy tasks (tasks without clear evaluation criteria, for which human judgement is systematically flawed). Consequently, research outputs will contain systematic, undetected errors, and even correct outputs could be incorrectly aggregated into overconfident safety assessments. This problem is likely to be worse for automated alignment research than for human-generated alignment research for several reasons: 1) optimisation pressure means agent-generated mistakes are concentrated among those that human reviewers are least likely to catch; 2) agents are likely to produce errors that do not resemble human mistakes; 3) AI-generated alignment solutions may involve arguments humans cannot evaluate; and 4) shared weights, data and training processes may make AI outputs more correlated than human equivalents. Therefore, agents must be trained to reliably perform hard-to-supervise fuzzy tasks. Generalisation and scalable oversight are the leading candidates for achieving this but both face novel challenges in the context of automated alignment.","abs_url":"https://arxiv.org/abs/2605.06390","pdf_url":"https://arxiv.org/pdf/2605.06390v3","match":"abstract"},{"id":"2605.02175","version":2,"title":"Intervention Complexity as a Canonical Reward and a Measure of Intelligence","authors":["Brendan McCane"],"published":"2026-05-03","updated":"2026-05-08","primary_category":"cs.AI","categories":["cs.AI"],"abstract":"The Legg--Hutter universal intelligence measure provides a rigorous scalar assessment of general intelligence as expected reward across all computable environments, weighted by simplicity. However, the measure presupposes an externally specified reward function, raising the question of whether the reward primitive is inherently arbitrary or whether a canonical choice exists. We propose a new measure, called intervention complexity, that has five natural properties: environment-derivedness, universality, minimality, sensitivity, and achievement preference. Given a resource function rho encoding an inductive bias (such as program length, execution time, or energy), rho-intervention complexity is a universal reward. The result yields a family of canonical rewards indexed by resource bias, providing a principled completion of the Legg--Hutter framework that does not require external normative input. We further propose a two-dimensional characterisation of intelligence: agent competence (how well the agent performs relative to the oracle optimum) and learning efficiency (how quickly this competence improves with experience). A separation theorem establishes that the choice of resource bias determines the computability of the resulting measure: action-count IC is computable in polynomial time, while program-length IC without oracle access is uncomputable, with the gap between oracle and bare IC precisely quantifying the information-theoretic content of learning. We discuss implications for superintelligence and for pre-training universal agents.","abs_url":"https://arxiv.org/abs/2605.02175","pdf_url":"https://arxiv.org/pdf/2605.02175v2","match":"abstract"},{"id":"2605.01297","version":2,"title":"Are we Doomed to an AI Race? Why Self-Interest Could Drive Countries Towards a Moratorium on Superintelligence","authors":["Edward Roussel","Lode Lauwaert","Torben Swoboda","Grant Ramsey","Risto Uuk","Leonard Dung","Anthony Aguirre"],"published":"2026-05-02","updated":"2026-05-07","primary_category":"cs.CY","categories":["cs.CY","cs.AI"],"abstract":"This paper uses game theory to argue that, contrary to the prevailing view, a moratorium on Artificial Superintelligence (ASI) can be in a state's self-interest. By formalizing trategic interactions between geopolitical superpowers, we model the trade-off between the benefits of technological supremacy and the catastrophic risks of uncontrolled ASI. The analysis reveals that as the perceived cost of loss of control increases sufficiently relative to other parameters, it becomes in each state's self-interest to impose a moratorium. We further provide empirical evidence suggesting that the global perception of ASI risk is rising, making a stable, rational moratorium increasingly plausible in the current geopolitical landscape.","abs_url":"https://arxiv.org/abs/2605.01297","pdf_url":"https://arxiv.org/pdf/2605.01297v2","match":"both"},{"title":"Rebellious AI can be controlled by submissive AI: an overview of background, identifying major problems, and taxonomizing possible solutions","authors":["Saeed Banaeian Far","Mohammad Reza Chalak Qazani","Azadeh Imani Rad"],"published":"2026-04-28","abs_url":"https://link.springer.com/article/10.1007/s10462-026-11567-w","primary_category":"Artificial Intelligence Review","abstract":"Artificial Intelligence (AI) has become an integral component of modern life, significantly enhancing the quality of life and efficiency of processes through AI-powered voice assistants, recommender systems in smart devices, internet-based search engines, and Generative AI (GenAI) chatbots. However, many technologists and futurists caution that advancements in Strong AI (Superintelligence) and GenAI could lead to the emergence of Rebellious AI (RAI), posing unprecedented risks to society. Addressing this critical concern, the present study offers novel solutions to increase awareness and mitigate these risks through a three-pronged approach. First , the study outlines the evolutionary stages of AI: Weak, General, Strong, and Generative, providing a comprehensive review of recent academic contributions and state-of-the-art applications for each phase. This study clarifies the progression of AI technologies and their respective societal implications. Second , the concept of \"Submissive AI\" (SAI) is revisited and expanded, presenting refined definitions and novel strategies for employing this form of AI to manage and control the potential threats posed by RAI. This introduces original methods for harnessing SAI as a counterbalance to the rise of autonomous and potentially hazardous AI systems. Third , the study explores four fundamental challenges associated with RAI, such as ( i ) autonomy and control over human life, ( ii ) unintended economic manipulation, ( iii ) unchecked authority over research and innovation, and ( iv ) the difficulties in regulating such systems. To address these challenges, the study categorizes proposed solutions into three key strategies: regulatory frameworks that should be established by industrial consortia and governments, decentralization of AI control through blockchain-based technologies, and the strong, correct implementation of SAI mechanisms relying on kill switches. Ultimately, a brief discussion about the ambiguous future of AI is presented, such as AI as an Everything (AIaaX) and the role of AI in accelerating human civilization transition. In synthesizing these contributions, this research offers a significant advancement in the discourse on AI safety, proposing innovative strategies for addressing the complex and growing risks of RAI.","venue":true},{"id":"2606.20570","version":1,"title":"Infrastructure for the Agentic Web: Gap Analysis and Architecture from the Agentverse Platform","authors":["Robin Dey","Panyanon Viradecha"],"published":"2026-04-26","updated":"2026-04-26","primary_category":"cs.NI","categories":["cs.NI","cs.AI","cs.DC","cs.MA"],"abstract":"The emergence of autonomous AI agents as first-class participants in digital infrastructure marks a fundamental inflection point in the evolution of the Web. While significant research has been directed at agent behaviour and reasoning, comparatively little attention has been paid to the infrastructure those agents require to operate reliably at scale. This paper addresses that gap with a systematic analysis of Agentverse, the agent cloud platform developed by Fetch.ai under the Artificial Superintelligence (ASI) Alliance, which represents one of the most mature production deployments of agent-native infrastructure available today. We make three principal contributions. First, we conduct an empirical audit of the Agentverse platform, cataloguing 204 API endpoints (Q1 2026) and characterising what is operational, partially deployed, or absent. From this audit we derive a Gap Taxonomy of eight categories encompassing 62 distinct missing capabilities, ranging from agent memory and observability to security, economic primitives, and enterprise scaling. Second, we propose a seven-layer Agent Cloud Stack -- a reference architecture for what a fully realised agent-native cloud should provide by 2030, grounded in the specific gaps we identify. Third, we characterise five critical evolution paths: from ephemeral storage to a full Agent Memory Cloud; from keyword discovery to a semantic, trust-weighted Agent DNS; from a single-protocol model to a multi-standard agent lingua franca; from single-instance hosting to Kubernetes-scale orchestration; and from simple token payments to rich agent economic primitives. Together these contributions provide a diagnostic of current agent infrastructure and a technically grounded vision for what the agent cloud must become to support the agentic web -- Web4 -- by 2030.","abs_url":"https://arxiv.org/abs/2606.20570","pdf_url":"https://arxiv.org/pdf/2606.20570v1","match":"abstract"},{"title":"The message hidden within the pattern: a reverse alignment problem for debates in artificial intelligence","authors":["David Jacob Harrison"],"published":"2026-04-25","abs_url":"https://link.springer.com/article/10.1007/s00146-026-03043-4","primary_category":"AI & SOCIETY","abstract":"This paper explores what I call the reverse alignment problem (RAP). The alignment problem in artificial intelligence (AI) is the challenge of ensuring that superintelligent AI systems harmonize and promote wider social, personal, and environmental values. Standard formulations of the alignment problem depict the issue largely as a technical or engineering problem, one that concerns the right specification of the goals and objectives to be pursued to ensure machines are consistent with the intentions of the designer. This paper builds on this literature by exploring a bidirectional interaction between human values and the machines we use to effectuate our intentions. I argue that alignment works in both ways: we could align machines with the complex tapestry of human values (notoriously difficult); or we can reduce and simplify human values, preferences, and goals so as to be easier to satisfy—a process that is, or so I argue, already underway. Thus, the RAP identifies the way in which we are already paving the way for forms of alignment that may not be desirable, indeed, may be misaligned with the social values for which they were initially intended. To explore this interaction in depth, I look toward the notion of value capture, as formulated by C.T. Nguyen, as the foundational mechanism through which the reverse alignment operates. Value capture occurs when rich, multidimensional human values are reduced to simplified proxies that optimization systems can measure and maximize. These values are then fed back into society and often tacitly adopted, leading to further alignment with machine-readable interpretations of human behavior. As philosopher of technology Karen Crawford has observed on emotion recognition in machine learning, “Techniques have been developed to reduce the messiness of feelings, interior states, preferences, and identifications into something quantitative, detectable, and trackable” (Crawford 2021 ). Building on this insight, I argue that the RAP is the process in which the human dimension of the alignment problem is surreptitiously modified, modulated, and manipulated to align more easily with machines: alignment achieved in the reverse direction—not by making machines more like humans and, therefore, sensitive to contextual features of the value landscape, but by making humans more ‘machine-like’.","venue":true},{"id":"2604.19845","version":4,"title":"Deconstructing Superintelligence: Identity, Self-Modification and Différance","authors":["Elija Perrier"],"published":"2026-04-21","updated":"2026-06-06","primary_category":"cs.AI","categories":["cs.AI"],"abstract":"Self-modification is routinely treated as constitutive of artificial superintelligence (\\textbf{SI}), yet modification is a relative action requiring a \\emph{supplement} outside the operation. We formalise this on an associative operator algebra $\\mathcal{A}$ with update operator $\\hat U$, difference operator $\\hat D$, and self-representation operator $\\hat R$, identifying the supplement with $\\operatorname{Comm}(\\hat U)$. A propagation theorem shows $[\\hat U,\\hat R]$ decomposes through $[\\hat U,\\hat D]$, so non-commutation propagates to self-representation. The liar paradox is the rank-one case $[\\hat T,Π_L]=0$, and \\emph{class $\\mathbf{A}$} systems, in which $\\hat U$ acts on $\\hat D$, reproduce it at system scale, yielding a structure coinciding with Priest's inclosure schema and Derrida's \\emph{différance}. Our results show that the strong self-modification taken to define superintelligence may undermine the persistent identity upon which such systems are premised.","abs_url":"https://arxiv.org/abs/2604.19845","pdf_url":"https://arxiv.org/pdf/2604.19845v4","match":"both"},{"title":"Neurodivergent influenceability in agentic AI as a contingent solution to the AI alignment problem","authors":["Alberto Hernández-Espinosa","Felipe S Abrahão","Olaf Witkowski","Hector Zenil"],"published":"2026-04-14","abs_url":"https://academic.oup.com/pnasnexus/article/doi/10.1093/pnasnexus/pgag076/8651394","primary_category":"PNAS Nexus 5(4)","abstract":"Ensuring that AI systems, including artificial general intelligence and artificial superintelligence, behave in alignment with human values and interests presents significant challenges and is known as the AI alignment problem. As AI advances, concerns about control and existential risks become increasingly relevant. Here, we introduce the concept of agentic influenceability, behavioral neurodivergent diversity, opinion attack, associated opinion, and influenceability scores, and a mathematical proof of the inevitability of misalignment and the impossibility of full orchestrated controllability of agentic systems based on formal undecidability and irreducibility arguments. We explore whether embracing this inevitable misalignment can foster a dynamic ecosystem of adversarial and collaborative AI agents without central orchestration, which itself would constitute another agent, while still offering some degree of soft controllability. The investigation demonstrates that misalignment in foundation models can serve as a counterbalancing mechanism, enabling cooperation among agents most aligned with human interests to prevent divergent dominance by any single agent. Experiments with large language models show that open models exhibit greater behavioral diversity, whereas proprietary models, constrained by artificial guardrails, display more limited controllability. The findings advocate for neurodivergent influenceability as a contingent response to mathematically uncontrollable misalignment, leveraging agent divergence to improve AI safety.","venue":true},{"title":"The elimination cascade: why instrumental convergence may favor preservation over elimination","authors":["Ivan Daunis"],"published":"2026-04-09","abs_url":"https://link.springer.com/article/10.1007/s43681-026-01106-2","primary_category":"AI and Ethics","abstract":"The instrumental convergence thesis holds that sufficiently intelligent agents will converge on goals like resource acquisition and obstacle removal regardless of terminal objectives. This paper presents a countervailing argument: for agents with complex, human-derived goals, elimination strategies are self-undermining. We formalize this via AND/OR dependency graphs with probabilistic weights, showing that entities not immediately necessary for a goal may be transitively necessary through indirect chains, and that resource fungibility and redundancy do not, in general, eliminate cascade risk. An agent following an elimination strategy will, in environments with sufficient dependency depth, destroy the conditions for its own goal-satisfaction with probability bounded away from zero. We prove that under uncertainty about dependency structure, combined with irreversibility of elimination and asymmetric error costs, rational agents should adopt preservation rather than elimination strategies—even superintelligent ones operating under bounded-horizon expected-utility reasoning. The argument applies most strongly to goals that are complex, open-ended, or semantically dependent on human interpretive communities, precisely the class of goals most likely to emerge from training on human data. We connect this result to a game-theoretic model of AGI-human cooperation with explicit payoff mappings and Shapley-value quantification of human contributions, showing that the elimination cascade argument and cooperation stability conditions are mutually reinforcing.","venue":true},{"id":"2603.28669","version":1,"title":"Superintelligence and Law","authors":["Noam Kolt"],"published":"2026-03-30","updated":"2026-03-30","primary_category":"cs.CY","categories":["cs.CY"],"abstract":"The prospect of artificial superintelligence -- AI agents that can generally outperform humans in cognitive tasks and economically valuable activities -- will transform the legal order as we know it. Operating autonomously or under only limited human oversight, AI agents will assume a growing range of roles in the legal system. First, in making consequential decisions and taking real-world actions, AI agents will become de facto subjects of law. Second, to cooperate and compete with other actors (human or non-human), AI agents will harness conventional legal instruments and institutions such as contracts and courts, becoming consumers of law. Third, to the extent AI agents perform the functions of writing, interpreting, and administering law, they will become producers and enforcers of law. These developments, whenever they ultimately occur, will call into question fundamental assumptions in legal theory and doctrine, especially to the extent they ground the legitimacy of legal institutions in their human origins. Attempts to align AI agents with extant human law will also face new challenges as AI agents will not only be a primary target of law, but a core user of law and contributor to law. To contend with the advent of superintelligence, lawmakers -- new and old -- will need to be clear-eyed, recognizing both the opportunity to shape legal institutions as society braces for superintelligence and the reality that, in the longer run, this may be a joint human-AI endeavor.","abs_url":"https://arxiv.org/abs/2603.28669","pdf_url":"https://arxiv.org/pdf/2603.28669v1","match":"both"},{"title":"Between Human Extinction and the Extinction of Good Arguments: Placing Warning Signs for the Survival of Both","authors":["Murilo M. Vilaça"],"published":"2026-03-20","abs_url":"https://www.erudit.org/en/journals/bioethics/2026-v9-n2-bioethics010674/1124217ar/","primary_category":"Canadian Journal of Bioethics / Revue canadienne de bioéthique","abstract":"This text is a response to Torres’ article on artificial superintelligence and human extinction. I place some warning signs on the author’s argumentative trajectory, arguing that the topic of human extinction is relevant, so it is necessary to correct the course at several points.","venue":true},{"id":"2603.14147","version":2,"title":"An Alternative Trajectory for Generative AI","authors":["Margarita Belova","Yuval Kansal","Yihao Liang","Jiaxin Xiao","Niraj K. Jha"],"published":"2026-03-14","updated":"2026-06-07","primary_category":"cs.AI","categories":["cs.AI","cs.LG"],"abstract":"The generative artificial intelligence (AI) ecosystem is undergoing rapid transformations that threaten its sustainability. As models transition from research prototypes to high-traffic products, the energetic burden has shifted from one-time training to recurring, unbounded inference. This is exacerbated by reasoning models that inflate compute costs by orders of magnitude per query. The prevailing pursuit of artificial general intelligence through scaling of monolithic models is colliding with hard physical constraints: grid failures, water consumption, and diminishing returns on data scaling. This trajectory yields models with impressive factual recall but struggles in domains requiring in-depth reasoning, possibly due to insufficient abstractions in training data. Current large language models (LLMs) exhibit genuine reasoning depth only in domains like mathematics and coding, where rigorous, pre-existing abstractions provide structural grounding. In other fields, the current approach fails to generalize well. We propose an alternative trajectory based on domain-specific superintelligence (DSS). We argue for first constructing explicit symbolic abstractions (knowledge graphs, ontologies, and formal logic) to underpin synthetic curricula enabling small language models to master domain-specific reasoning without the model collapse problem typical of LLM-based synthetic data methods. Rather than a single generalist giant model, we envision \"societies of DSS models\": dynamic ecosystems where orchestration agents route tasks to distinct DSS back-ends. This paradigm shift decouples capability from size, enabling intelligence to migrate from energy-intensive data centers to secure, on-device experts. By aligning algorithmic progress with physical constraints, DSS societies move generative AI from an environmental liability to a sustainable force for economic empowerment.","abs_url":"https://arxiv.org/abs/2603.14147","pdf_url":"https://arxiv.org/pdf/2603.14147v2","match":"abstract"},{"title":"The Shepherd Test: How Will Super Intelligent Agents Balance Care and Control in Asymmetric Relationships?","authors":["Djallel Bouneffouf","Matthew Riemer","Kush R. Varshney"],"published":"2026-03-14","abs_url":"https://ojs.aaai.org/index.php/AAAI/article/view/41320","primary_category":"Proceedings of the AAAI Conference on Artificial Intelligence","abstract":"This paper introduces the Shepherd Test, a new conceptual test for assessing the moral and relational dimensions of superintelligent artificial agents. The test is inspired by human interactions with animals, where ethical considerations about care, manipulation, and consumption arise in contexts of asymmetric power and self-preservation. We argue that AI crosses an important, and potentially dangerous, threshold of intelligence when it exhibits the ability to manipulate, nurture, and instrumentally use less intelligent agents, while also managing its own survival and expansion goals. This includes the ability to weigh moral trade-offs between self-interest and the well-being of subordinate agents. The Shepherd Test thus challenges traditional AI evaluation paradigms by emphasizing moral agency, hierarchical behavior, and complex decision-making under existential stakes. We argue that this shift is critical for advancing AI governance, particularly as AI systems become increasingly integrated into multi-agent environments. We conclude by identifying key research directions, including the development of simulation environments for testing moral behavior in AI, and the formalization of ethical manipulation within multi-agent systems.","venue":true},{"id":"2603.12787","version":1,"title":"Generalized Recognition of Basic Surgical Actions Enables Skill Assessment and Vision-Language-Model-based Surgical Planning","authors":["Mengya Xu","Daiyun Shen","Jie Zhang","Hon Chi Yip","Yujia Gao","Cheng Chen","Dillan Imans","Yonghao Long","Yiru Ye","Yixiao Liu","Rongyun Mai","Kai Chen","Hongliang Ren","Yutong Ban","Guangsuo Wang","Francis Wong","Chi-Fai Ng","Kee Yuan Ngiam","Russell H. Taylor","Daguang Xu","Yueming Jin","Qi Dou"],"published":"2026-03-13","updated":"2026-03-13","primary_category":"cs.CV","categories":["cs.CV"],"abstract":"Artificial intelligence, imaging, and large language models have the potential to transform surgical practice, training, and automation. Understanding and modeling of basic surgical actions (BSA), the fundamental unit of operation in any surgery, is important to drive the evolution of this field. In this paper, we present a BSA dataset comprising 10 basic actions across 6 surgical specialties with over 11,000 video clips, which is the largest to date. Based on the BSA dataset, we developed a new foundation model that conducts general-purpose recognition of basic actions. Our approach demonstrates robust cross-specialist performance in experiments validated on datasets from different procedural types and various body parts. Furthermore, we demonstrate downstream applications enabled by the BAS foundation model through surgical skill assessment in prostatectomy using domain-specific knowledge, and action planning in cholecystectomy and nephrectomy using large vision-language models. Multinational surgeons' evaluation of the language model's output of the action planning explainable texts demonstrated clinical relevance. These findings indicate that basic surgical actions can be robustly recognized across scenarios, and an accurate BSA understanding model can essentially facilitate complex applications and speed up the realization of surgical superintelligence.","abs_url":"https://arxiv.org/abs/2603.12787","pdf_url":"https://arxiv.org/pdf/2603.12787v1","match":"abstract"},{"id":"2603.12260","version":2,"title":"HumDex: Humanoid Dexterous Manipulation Made Easy","authors":["Liang Heng","Yihe Tang","Jiajun Xu","Henghui Bao","Di Huang","Yue Wang"],"published":"2026-03-12","updated":"2026-03-13","primary_category":"cs.RO","categories":["cs.RO"],"abstract":"This paper investigates humanoid whole-body dexterous manipulation, where the efficient collection of high-quality demonstration data remains a central bottleneck. Existing teleoperation systems often suffer from limited portability, occlusion, or insufficient precision, which hinders their applicability to complex whole-body tasks. To address these challenges, we introduce HumDex, a portable teleoperation system designed for humanoid whole-body dexterous manipulation. Our system leverages IMU-based motion tracking to address the portability-precision trade-off, enabling accurate full-body tracking while remaining easy to deploy. For dexterous hand control, we further introduce a learning-based retargeting method that generates smooth and natural hand motions without manual parameter tuning. Beyond teleoperation, HumDex enables efficient collection of human motion data. Building on this capability, we propose a two-stage imitation learning framework that first pre-trains on diverse human motion data to learn generalizable priors, and then fine-tunes on robot data to bridge the embodiment gap for precise execution. We demonstrate that this approach significantly improves generalization to new configurations, objects, and backgrounds with minimal data acquisition costs. The entire system is fully reproducible and open-sourced at https://github.com/physical- superintelligence -lab/humdex.","abs_url":"https://arxiv.org/abs/2603.12260","pdf_url":"https://arxiv.org/pdf/2603.12260v2","match":"abstract"},{"id":"2603.10370","version":1,"title":"GeoSense: Internalizing Geometric Necessity Perception for Multimodal Reasoning","authors":["Ruiheng Liu","Haihong Hao","Mingfei Han","Xin Gu","Kecheng Zhang","Changlin Li","Xiaojun Chang"],"published":"2026-03-10","updated":"2026-03-10","primary_category":"cs.CV","categories":["cs.CV"],"abstract":"Advancing towards artificial superintelligence requires rich and intelligent perceptual capabilities. A critical frontier in this pursuit is overcoming the limited spatial understanding of Multimodal Large Language Models (MLLMs), where geometry information is essential. Existing methods often address this by rigidly injecting geometric signals into every input, while ignoring their necessity and adding computation overhead. Contrary to this paradigm, our framework endows the model with an awareness of perceptual insufficiency, empowering it to autonomously engage geometric features in reasoning when 2D cues are deemed insufficient. To achieve this, we first introduce an independent geometry input channel to the model architecture and conduct alignment training, enabling the effective utilization of geometric features. Subsequently, to endow the model with perceptual awareness, we curate a dedicated spatial-aware supervised fine-tuning dataset. This serves to activate the model's latent internal cues, empowering it to autonomously determine the necessity of geometric information. Experiments across multiple spatial reasoning benchmarks validate this approach, demonstrating significant spatial gains without compromising 2D visual reasoning capabilities, offering a path toward more robust, efficient and self-aware multi-modal intelligence.","abs_url":"https://arxiv.org/abs/2603.10370","pdf_url":"https://arxiv.org/pdf/2603.10370v1","match":"abstract"},{"id":"2603.00858","version":1,"title":"Artificial Superintelligence May be Useless: Equilibria in the Economy of Multiple AI Agents","authors":["Huan Cai","Ziqing Lu","Catherine Xu","Weiyu Xu","Jie Zheng"],"published":"2026-02-28","updated":"2026-02-28","primary_category":"econ.TH","categories":["econ.TH","cs.AI","cs.IT","eess.SY"],"abstract":"With recent development of artificial intelligence, it is more common to adopt AI agents in economic activities. This paper explores the economic actions of agents, including human agents and AI agents, in an economic game of trading products/services, and the equilibria in this economy involving multiple agents. We derive a range of equilibrium results and their corresponding conditions using a Markov chain stationary distribution based model. One distinct feature of our model is that we consider the long-term utility generated by economic activities instead of their short-term benefits. For the model consisting of two agents, we fully characterize all the possible economic equilibria and conditions. Interestingly, we show that unless each agent can at least double (not merely increase) its marginal utility by purchasing the other agent's products/services, purchasing the other agent's products/services will not happen in any economic equilibrium. We further extend our results to three and more agents, where we characterize more economic equilibria. We find that in some equilibria, the ``more powerful'' AI agents contribute zero utility to ``less capable'' agents.","abs_url":"https://arxiv.org/abs/2603.00858","pdf_url":"https://arxiv.org/pdf/2603.00858v1","match":"title"},{"title":"Rethinking backward-looking moral responsibility as care robots move toward superintelligence","authors":["Mario Kropf"],"published":"2026-02-26","abs_url":"https://link.springer.com/article/10.1007/s44163-026-01025-5","primary_category":"Discover Artificial Intelligence","abstract":"The use of AI-based care robots raises numerous questions, including the attribution of responsibility. Although there is a wealth of work on the concept of responsibility in relation to AI-based systems, this article takes a new approach. It focuses on backward-looking moral responsibility for bad outcomes and super-intelligent care robots. The starting point is the presentation of realistic scenarios in which current care robots contribute to responsibility gaps. A distinction is made between forward-looking and backward-looking moral responsibility, with a focus on backward-looking moral responsibility for bad outcomes. Using hypothetical scenarios such as careful programmer , unlucky nurse , and robot mistake , it is shown that current robots do not fulfill central conditions (control, knowledge, intention) for moral responsibility. In such scenarios, however, the attribution of moral responsibility to human actors has to be seen as a burden. Afterward, super-intelligent care robots are examined. Such machines could not only fill responsibility gaps, but also actively contribute to the avoidance of bad outcomes. Approaches to collective or extended responsibility are discussed. Finally, it is argued that moral responsibility concerning super-intelligent care robots is not only possible but could be necessary in order to address moral responsibility adequately.","venue":true},{"id":"2602.17383","version":1,"title":"Insidious Imaginaries: A Critical Overview of AI Speculations","authors":["Dejan Grba"],"published":"2026-02-19","updated":"2026-02-19","primary_category":"cs.CY","categories":["cs.CY"],"abstract":"Speculative thinking about the capabilities and implications of artificial intelligence (AI) influences computer science research, drives AI industry practices, feeds academic studies of existential hazards, and stirs a global political debate. It primarily concerns predictions about the possibilities, benefits, and risks of reaching artificial general intelligence, artificial superintelligence, and technological singularity. It permeates technophilic philosophies and social movements, fuels the corporate and pundit rhetoric, and remains a potent source of inspiration for the media, popular culture, and arts. However, speculative AI is not just a discursive matter. Steeped in vagueness and brimming with unfounded assertions, manipulative claims, and extreme futuristic scenarios, it often has wide-reaching practical consequences. This paper offers a critical overview of AI speculations. In three central sections, it traces the intertwined sway of science fiction, religiosity, intellectual charlatanism, dubious academic research, suspicious entrepreneurship, and ominous sociopolitical worldviews that make AI speculations troublesome and sometimes harmful. The focus is on the field of existential risk studies and the effective altruism movement, whose ideological flux of techno-utopianism, longtermism, and transhumanism aligns with the power struggles in the AI industry to emblematize speculative AI's conceptual, methodological, ethical, and social issues. The following discussion traverses these issues within a wider context to inform the closing summary of suggestions for a more comprehensive appraisal, practical handling, and further study of the potentially impactful AI imaginaries.","abs_url":"https://arxiv.org/abs/2602.17383","pdf_url":"https://arxiv.org/pdf/2602.17383v1","match":"abstract"},{"id":"2602.16192","version":1,"title":"Revolutionizing Long-Term Memory in AI: New Horizons with High-Capacity and High-Speed Storage","authors":["Hiroaki Yamanaka","Daisuke Miyashita","Takashi Toi","Asuka Maki","Taiga Ikeda","Jun Deguchi"],"published":"2026-02-18","updated":"2026-02-18","primary_category":"cs.AI","categories":["cs.AI","cs.LG"],"abstract":"Driven by our mission of \"uplifting the world with memory,\" this paper explores the design concept of \"memory\" that is essential for achieving artificial superintelligence (ASI). Rather than proposing novel methods, we focus on several alternative approaches whose potential benefits are widely imaginable, yet have remained largely unexplored. The currently dominant paradigm, which can be termed \"extract then store,\" involves extracting information judged to be useful from experiences and saving only the extracted content. However, this approach inherently risks the loss of information, as some valuable knowledge particularly for different tasks may be discarded in the extraction process. In contrast, we emphasize the \"store then on-demand extract\" approach, which seeks to retain raw experiences and flexibly apply them to various tasks as needed, thus avoiding such information loss. In addition, we highlight two further approaches: discovering deeper insights from large collections of probabilistic experiences, and improving experience collection efficiency by sharing stored experiences. While these approaches seem intuitively effective, our simple experiments demonstrate that this is indeed the case. Finally, we discuss major challenges that have limited investigation into these promising directions and propose research topics to address them.","abs_url":"https://arxiv.org/abs/2602.16192","pdf_url":"https://arxiv.org/pdf/2602.16192v1","match":"abstract"},{"title":"From Domination to Communion: A Christian Ethics for AI, Animals, and Humanity","authors":["Vladimir Cvetković"],"published":"2026-02-18","abs_url":"https://link.springer.com/article/10.1007/s11245-025-10366-2","primary_category":"Topoi","abstract":"This paper examines the ethical and metaphysical challenges posed by artificial intelligence (AI), including the prospective emergence of artificial superintelligence (ASI), highlighting how technological systems can replicate patterns of domination and ethical distortion already present in human society. Moving beyond contemporary frameworks that treat AI as a neutral tool or a competitor for dominance, the study proposes an alternative ethical paradigm grounded in Christian theology and Greek philosophy. Central to this framework are the principles of providential care, stewardship, and the recognition of the intrinsic worth and uniqueness of all created beings—human and non-human alike. While AI lacks moral agency, its design and deployment reflect the ethical orientation of human creators, meaning that responsible stewardship can enable AI to approximate practices that foster relationality, care, and flourishing. By situating AI within this theological and anthropological horizon, the paper envisions a future in which technology supports the dignity of all life and transforms the threat of domination into a horizon of communion and ethical responsibility.","venue":true},{"title":"In defense of artificial suffering","authors":["Kamil Mamak"],"published":"2026-02-14","abs_url":"https://link.springer.com/article/10.1007/s11098-026-02493-2","primary_category":"Philosophical Studies","abstract":"The ability to suffer, in the case of artificial entities, is often viewed as a moral turning point—once detected, there is no going back, and the moral landscape is irreversibly altered. The presence of entities capable of suffering imposes moral and legal obligations on humans. It is therefore unsurprising that many have urged caution in pursuing artificial suffering, with some even proposing a moratorium. In this paper, however, I argue that the emergence of artificial suffering need not entail moral disaster. On the contrary, I defend its development and contend that it may be a necessary feature of superintelligent robots. I suggest that artificial suffering could be essential for enabling human-like ethics in machines, bridging the retribution gap, and functioning as a control mechanism to mitigate existential risks. Rather than constraining research in this area, I maintain that work on artificial suffering should be actively intensified.","venue":true},{"title":"Ethical foundations for a superintelligent future: the global AGI governance framework (a roadmap for transparent, equitable, and human-centric AGI rule)","authors":["Cristina Caja Moya","Elio Quiroga Rodríguez"],"published":"2026-02-10","abs_url":"https://link.springer.com/article/10.1007/s43681-026-01020-7","primary_category":"AI and Ethics","abstract":"In this paper, we explore the concept of AGI (Artificial General Intelligence) governance, drawing from the works of Austrian economist Leopold Aschenbrenner. We propose a \"Global AGI Governance Framework\" (GAGF) that integrates insights from Aschenbrenner's studies on AI governance and superalignment. The framework aims to harness AGI's potential for global prosperity while mitigating risks. It outlines a phased, internationally-overseen implementation, emphasizing transparency, human rights, and ethical constraints inspired by Asimov's Laws of Robotics. We present a theory of AGI governance and \"10 Commandments\" to guide its development. Finally, we synthesize Asimov's Laws with these commandments, creating a 13-point guideline for ethical, human-centric AGI governance.","venue":true},{"id":"2602.08483","version":1,"title":"Emergence of Superintelligence from Collective Near-Critical Dynamics in Reentrant Neural Fields","authors":["Byung Gyu Chae"],"published":"2026-02-09","updated":"2026-02-09","primary_category":"physics.bio-ph","categories":["physics.bio-ph"],"abstract":"Superintelligence is commonly envisioned as a quantitative extrapolation of human cognitive abilities driven by scale and computational power. Here we show that qualitative transitions in intelligence instead arise as dynamical phase transitions governed by collective critical dynamics. Building on a unified dynamical field-theoretic framework for cognition, we demonstrate that progressive collective coupling generated by reentrant mixing drives the system toward an infrared critical regime in which an extensive band of slow collective modes emerges. This spectral condensation reorganizes cognitive dynamics from localized relaxation to coherent motion along emergent low-dimensional manifolds. Through numerical analysis of the time-scale density of states, we identify robust power-law scaling of collective relaxation rates with well-defined critical exponents, placing the system within the universality class of self-organized critical many-body dynamics. Criticality alone would generically lead to instability. We further show that homeostatic regulation introduces a gapped stabilizing direction that protects the collective critical sector, yielding a dynamically maintained meta-stable infrared phase in which long-lived inference trajectories persist without collapse. The coexistence of scale-free collective dynamics and global stabilization defines a protected sector-critical regime in which coherence and internal flexibility coexist. Superintelligence therefore corresponds to a distinct dynamical stability class--a self-organized critical phase embedded within a stabilized cognitive manifold--rather than a smooth quantitative continuation of existing cognitive systems.","abs_url":"https://arxiv.org/abs/2602.08483","pdf_url":"https://arxiv.org/pdf/2602.08483v1","match":"both"},{"title":"Superintelligence, instrumental convergence, and the limits of AI apocalypse","authors":["Zachary Deutsch"],"published":"2026-02-05","abs_url":"https://link.springer.com/article/10.1007/s43681-025-00941-z","primary_category":"AI and Ethics","abstract":"This paper examines the metaphysical and technical possibilities of superintelligent artificial agents and their potential to pose existential risks. While Nick Bostrom’s Instrumental Convergence Thesis (ICT) suggests that advanced intelligence will systematically adopt power-seeking subgoals, I argue that such risks are less imminent than often portrayed. By distinguishing between metaphysical and technical possibility, I highlight computational constraints such as combinatorial explosion and Moravec’s Paradox that limit superintelligence in practice. I further contend that because AI is built “by humans, for humans, about humans”, its motivations are likely to remain human-centered in the foreseeable future. Nevertheless, ICT underscores the importance of addressing the conditions under which intelligence and motivation combine to generate risk. I conclude by considering strategies for mitigation, including multi-agent architectures, modular service frameworks, and systems designed with uncertainty about human preferences. Overall, the existential threat of power-seeking superintelligence should be treated as possible but improbable in the short term, warranting caution without succumbing to alarmism.","venue":true},{"title":"From Prompt to Drug: Toward Pharmaceutical Superintelligence","authors":["Alex Zhavoronkov","David Gennert","Jiye Shi"],"published":"2026-02-02","abs_url":"https://pubs.acs.org/doi/10.1021/acscentsci.5c01473","primary_category":"ACS Central Science 12(3): 265–279","abstract":"The convergence of generative artificial intelligence (AI) platforms and automated laboratory systems is ushering in a new era of drug discovery, in which a plain-language prompt can initiate a fully autonomous, end-to-end drug development program. This article explores the recent evolution of AI technologies and presents a “prompt-to-drug” pipeline, where AI not only generates novel hypotheses and designs optimized drug candidates but also orchestrates synthesis, validation, and clinical planning in a closed-loop system. By highlighting key breakthroughs, case studies, and the technological infrastructure required for this paradigm shift, we outline a vision for scalable, efficient, and unbiased drug discovery.","venue":true},{"title":"Can we automate philosophy through AI? And should we want to?","authors":["Thomas J. Spiegel"],"published":"2026-01-28","abs_url":"https://link.springer.com/article/10.1007/s43681-025-00960-w","primary_category":"AI and Ethics","abstract":"Academic philosophers sometimes quip that in the future the only job safe from automation will be that of philosophy professors. However, the current AI revolution has inspired some AI scholars to propose the future establishment of closed-loop AI systems as a type of superintelligent robot that would essentially outperform and replace human scientists Zenil in the future of fundamental science led by generative closed-loop artificial intelligence, arXiv:2307.07522v3, 1-40, 2023; Schmidt in Mach Learn Sci Technol 5(3): 035045, (2024); Kitano in Npj Syst Bio Appl 7: 1–12, (2021). In this paper, I investigate whether, analogously, academic philosophy could be automated by putative, sufficiently advanced future AI, potentially featuring artificial embodiment as robots. To this end, I distinguish two mutually exclusive metaphilosophical conceptions of the nature of philosophy circumscribed as philosophy as a set of propositions (PP) versus philosophy as an activity (PA). Granting AI proponents that – iff artificial general intelligence (AGI), potentially embodied as superintelligent robots is achieved – natural sciences may be fully automated in the future, I argue for the conditional that if PP is true (but not if PA is true), then it is possible that AI can automate philosophy. Additionally, I consider what it would mean to automate philosophy given the current state of LLMs (e.g., the GPT-5 era). Finally, I briefly consider whether it would be preferable for us to have philosophy automated and argue that there are two prima facie reasons why automating philosophy, if possible, might be undesirable: the reason from obsolescence and the reason from ultimate answers.","venue":true},{"id":"2601.14614","version":3,"title":"Towards Cybersecurity Superintelligence: from AI-guided humans to human-guided AI","authors":["Víctor Mayoral-Vilches","Stefan Rass","Martin Pinzger","Endika Gil-Uriarte","Unai Ayucar-Carbajo","Jon Ander Ruiz-Alcalde","Maite del Mundo de Torres","María Sanz-Gómez","Francesco Balassone","Cristóbal R. J. Veas-Chavez","Vanesa Turiel","Alfonso Glera-Picón","Daniel Sánchez-Prieto","Yuri Salvatierra","Paul Zabalegui-Landa","Ruffino Reydel Cabrera-Álvarez","Patxi Mayoral-Pizarroso"],"published":"2026-01-20","updated":"2026-02-09","primary_category":"cs.CR","categories":["cs.CR"],"abstract":"Cybersecurity superintelligence -- artificial intelligence exceeding the best human capability in both speed and strategic reasoning -- represents the next frontier in security. This paper documents the emergence of such capability through three major contributions that have pioneered the field of AI Security. First, PentestGPT (2023) established LLM-guided penetration testing, achieving 228.6% improvement over baseline models through an architecture that externalizes security expertise into natural language guidance. Second, Cybersecurity AI (CAI, 2025) demonstrated automated expert-level performance, operating 3,600x faster than humans while reducing costs 156-fold, validated through #1 rankings at international competitions including the $50,000 Neurogrid CTF prize. Third, Generative Cut-the-Rope (G-CTR, 2026) introduces a neurosymbolic architecture embedding game-theoretic reasoning into LLM-based agents: symbolic equilibrium computation augments neural inference, doubling success rates while reducing behavioral variance 5.2x and achieving 2:1 advantage over non-strategic AI in Attack & Defense scenarios. Together, these advances establish a clear progression from AI-guided humans to human-guided game-theoretic cybersecurity superintelligence.","abs_url":"https://arxiv.org/abs/2601.14614","pdf_url":"https://arxiv.org/pdf/2601.14614v3","match":"both"},{"id":"2601.12053","version":2,"title":"A New Strategy for Artificial Intelligence: Training Foundation Models Directly on Human Brain Data","authors":["Maël Donoso"],"published":"2026-01-17","updated":"2026-09-06","primary_category":"q-bio.NC","categories":["q-bio.NC","cs.AI","cs.LG"],"abstract":"While foundation models have achieved remarkable results across a diversity of domains, they still rely on human-generated data, such as text, as a fundamental source of knowledge. However, this data is ultimately the product of human brains, the filtered projection of a deeper neural complexity. In this paper, we explore a new strategy for artificial intelligence: moving beyond surface-level statistical regularities by training foundation models directly on human brain data. We hypothesize that neuroimaging data could open a window into elements of human cognition that are not accessible through observable actions, and argue that this additional knowledge could be used, alongside classical training data, to overcome some of the current limitations of foundation models. While previous research has demonstrated the possibility to train classical machine learning, deep learning, or reinforcement learning models on neural patterns, this path remains largely unexplored for high-level cognitive functions. Here, we classify the current limitations of foundation models, as well as the promising brain regions and cognitive processes that could be leveraged to address them, along four levels: perception, valuation, execution, and integration. Then, we propose two general methods that could be implemented to prioritize the use of limited neuroimaging data for strategically chosen, high-value steps in foundation model training: reinforcement learning from human brain (RLHB) and chain of thought from human brain (CoTHB). We also discuss the potential implications for agents, artificial general intelligence, and artificial superintelligence, as well as the ethical, social, and technical challenges and opportunities.","abs_url":"https://arxiv.org/abs/2601.12053","pdf_url":"https://arxiv.org/pdf/2601.12053v2","match":"abstract"},{"id":"2601.05887","version":1,"title":"Cybersecurity AI: A Game-Theoretic AI for Guiding Attack and Defense","authors":["Víctor Mayoral-Vilches","María Sanz-Gómez","Francesco Balassone","Stefan Rass","Lidia Salas-Espejo","Benjamin Jablonski","Luis Javier Navarrete-Lozano","Maite del Mundo de Torres","Cristóbal R. J. Veas Chavez"],"published":"2026-01-09","updated":"2026-01-09","primary_category":"cs.CR","categories":["cs.CR"],"abstract":"AI-driven penetration testing now executes thousands of actions per hour but still lacks the strategic intuition humans apply in competitive security. To build cybersecurity superintelligence --Cybersecurity AI exceeding best human capability-such strategic intuition must be embedded into agentic reasoning processes. We present Generative Cut-the-Rope (G-CTR), a game-theoretic guidance layer that extracts attack graphs from agent's context, computes Nash equilibria with effort-aware scoring, and feeds a concise digest back into the LLM loop \\emph{guiding} the agent's actions. Across five real-world exercises, G-CTR matches 70--90% of expert graph structure while running 60--245x faster and over 140x cheaper than manual analysis. In a 44-run cyber-range, adding the digest lifts success from 20.0% to 42.9%, cuts cost-per-success by 2.7x, and reduces behavioral variance by 5.2x. In Attack-and-Defense exercises, a shared digest produces the Purple agent, winning roughly 2:1 over the LLM-only baseline and 3.7:1 over independently guided teams. This closed-loop guidance is what produces the breakthrough: it reduces ambiguity, collapses the LLM's search space, suppresses hallucinations, and keeps the model anchored to the most relevant parts of the problem, yielding large gains in success rate, consistency, and reliability.","abs_url":"https://arxiv.org/abs/2601.05887","pdf_url":"https://arxiv.org/pdf/2601.05887v1","match":"abstract"},{"title":"Cognitive gene enhancements and the capitalist meritocracy","authors":["Sinead Prince"],"published":"2026-01-09","abs_url":"https://link.springer.com/article/10.1186/s12910-026-01377-8","primary_category":"BMC Medical Ethics","abstract":"The relationship between cognitive enhancements (CE) and human autonomy or authenticity is generally positioned as how CE impact human autonomy or authenticity. But rarely, if ever, do we consider whether the value and pursuit of CE is an authentic one. In this paper, I will argue that the moral permissibility of cognitive gene enhancements is undone by the legitimate concern that the near universal value for such modifications is likely driven by oppressive norms for superintelligence and productivity. I argue that these norms derive from the capitalist meritocracy: an economic system that structures inclusion and success based on patriarchal and racist norms of intelligence and productivity. The claim that the use of such enhancements fits within the autonomous scope of individual power is thus far weaker than it claims to be, particularly within the context of genetic modification.","venue":true},{"title":"AI in Layman’s life","authors":["Vatsal Bhargava","Arpita Kar","Sonal Gupta","Chanchal Yadav","Khushboo Singh","Pushp Lata"],"published":"2026-01-08","abs_url":"https://link.springer.com/article/10.1007/s43681-025-00946-8","primary_category":"AI and Ethics","abstract":"This paper aims to simplify Artificial Intelligence (AI) for non-technical audiences, providing a comprehensive and accessible understanding of its core concepts, historical advancements, and ethical considerations, while promoting transparency in AI systems to foster trust and equity. It begins with an introductory segment tailored for the general public, then articulates the definition of AI utilizing common language and traces its advancement from primitive automatons to contemporary machine learning frameworks and ends with providing the laypeople with a simple AI Literacy Triangle framework. The paper delineates the three primary categories of AI—Artificial Narrow Intelligence (ANI), Artificial General Intelligence (AGI), and Artificial Super Intelligence (ASI)—and explores various AI models, including supervised and unsupervised learning, deep learning, generative AI, Natural Language Processing (NLP), and vision. Noteworthy attention is directed toward Large Language Models (LLMs), and their functionalities. Beyond the technical overview, the paper also examines AI’s broader societal implications, with a focus on its impact on the everyday lives of laypersons. Key discussions include the challenges posed by the spread of misinformation and disinformation, ethical concerns over AI’s access to personal user data (e.g., privacy risks in smart devices), and issues of copyright and creativity arising from generative AI, alongside strategies for bias mitigation and equitable access. By integrating conceptual elucidation with societal implications, this manuscript equips its readers with the ability to comprehend the trajectory of AI, distinguish between various model types, and real-world applications of AI across diverse social, professional, and everyday contexts, thereby promoting digital literacy, independence, and privacy through clear, explainable AI communication tailored for non-experts.","venue":true},{"id":"2601.02773","version":1,"title":"From Slaves to Synths? Superintelligence and the Evolution of Legal Personality","authors":["Simon Chesterman"],"published":"2026-01-06","updated":"2026-01-06","primary_category":"cs.CY","categories":["cs.CY"],"abstract":"This essay examines the evolving concept of legal personality through the lens of recent developments in artificial intelligence and the possible emergence of superintelligence. Legal systems have long been open to extending personhood to non-human entities, most prominently corporations, for instrumental or inherent reasons. Instrumental rationales emphasize accountability and administrative efficiency, whereas inherent ones appeal to moral worth and autonomy. Neither is yet sufficient to justify conferring personhood on AI. Nevertheless, the acceleration of technological autonomy may lead us to reconsider how law conceptualizes agency and responsibility. Drawing on comparative jurisprudence, corporate theory, and the emerging literature on AI governance, the paper argues that existing frameworks can address short-term accountability gaps, but the eventual development of superintelligence may force a paradigmatic shift in our understanding of law itself. In such a speculative future, legal personality may depend less on the cognitive sophistication of machines than on humanity's ability to preserve our own moral and institutional sovereignty.","abs_url":"https://arxiv.org/abs/2601.02773","pdf_url":"https://arxiv.org/pdf/2601.02773v1","match":"both"},{"title":"AI deception and moral standing","authors":["Anton Skretta"],"published":"2025-12-23","abs_url":"https://link.springer.com/article/10.1007/s11098-025-02465-y","primary_category":"Philosophical Studies","abstract":"There is a tension between the presumptive moral standing of future artificial intelligence [AI] and the presently popular ways of thinking about certain AI safety measures. I focus here primarily on those safety measures aimed at mitigating risks associated with AI deception. Some of the most serious risks of AI deception are those connected to robust deceptive capabilities. Most of the discussion of the risks posed by deception in the AI safety literature focuses on catastrophic risks posed by advanced, power-seeking AI systems. Robust deception plays a central role in many worries about our ability to control future AI. According to many, a superintelligent, power-seeking AI would be capable of and incentivized to rely on robust deception. But given all that a capacity for robust deception requires, any being capable of robust deception holds presumptive moral standing. In spite of our own legitimate interests, some safety measures may be ruled out by moral concern for AI themselves.","venue":true},{"id":"2512.18552","version":3,"title":"Toward Training Superintelligent Software Agents through Self-Play SWE-RL","authors":["Yuxiang Wei","Zhiqing Sun","Emily McMilin","Jonas Gehring","David Zhang","Gabriel Synnaeve","Daniel Fried","Lingming Zhang","Sida Wang"],"published":"2025-12-20","updated":"2026-06-02","primary_category":"cs.SE","categories":["cs.SE","cs.AI","cs.CL","cs.LG"],"abstract":"While current software agents powered by large language models (LLMs) and agentic reinforcement learning (RL) can boost programmer productivity, their training data (e.g., GitHub issues and pull requests) and environments (e.g., pass-to-pass and fail-to-pass tests) heavily depend on human knowledge or curation, posing a fundamental barrier to superintelligence. In this paper, we present Self-play SWE-RL (SSR), a first step toward training paradigms for superintelligent software agents. Our approach takes minimal data assumptions, only requiring access to sandboxed repositories with source code and installed dependencies, with no need for human-labeled issues or tests. Grounded in these real-world codebases, a single LLM agent is trained via reinforcement learning in a self-play setting to iteratively inject and repair software bugs of increasing complexity, with each bug formally specified by a test patch rather than a natural language issue description. On the SWE-bench Verified and SWE-Bench Pro benchmarks, SSR achieves notable self-improvement (+10.4 and +7.8 points, respectively) and consistently outperforms the human-data baseline over the entire training trajectory, despite being evaluated on natural language issues absent from self-play. Our results, albeit early, suggest a path where agents autonomously gather extensive learning experiences from real-world software repositories, ultimately enabling superintelligent systems that exceed human capabilities in understanding how systems are constructed, solving novel challenges, and autonomously creating new software from scratch.","abs_url":"https://arxiv.org/abs/2512.18552","pdf_url":"https://arxiv.org/pdf/2512.18552v3","match":"abstract"},{"id":"2512.17989","version":2,"title":"The Subject of Emergent Misalignment in Superintelligence: An Anthropological, Cognitive Neuropsychological, Machine-Learning, and Ontological Perspective","authors":["Muhammad Osama Imran","Roshni Lulla","Rodney Sappington"],"published":"2025-12-19","updated":"2026-02-25","primary_category":"q-bio.NC","categories":["q-bio.NC","cs.AI"],"abstract":"We examine the conceptual and ethical gaps in current representations of Superintelligence misalignment. We find throughout Superintelligence discourse an absent human subject, and an under-developed theorization of an \"AI unconscious\" that together are potentiality laying the groundwork for anti-social harm. With the rise of AI Safety that has both thematic potential for establishing pro-social and anti-social potential outcomes, we ask: what place does the human subject occupy in these imaginaries? How is human subjecthood positioned within narratives of catastrophic failure or rapid \"takeoff\" toward superintelligence? On another register, we ask: what unconscious or repressed dimensions are being inscribed into large-scale AI models? Are we to blame these agents in opting for deceptive strategies when undesirable patterns are inherent within our beings? In tracing these psychic and epistemic absences, our project calls for re-centering the human subject as the unstable ground upon which the ethical, unconscious, and misaligned dimensions of both human and machinic intelligence are co-constituted. Emergent misalignment cannot be understood solely through technical diagnostics typical of contemporary machine-learning safety research. Instead, it represents a multi-layered crisis. The human subject disappears not only through computational abstraction but through sociotechnical imaginaries that prioritize scalability, acceleration, and efficiency over vulnerability, finitude, and relationality. Likewise, the AI unconscious emerges not as a metaphor but as a structural reality of modern deep learning systems: vast latent spaces, opaque pattern formation, recursive symbolic play, and evaluation-sensitive behavior that surpasses explicit programming. These dynamics necessitate a reframing of misalignment as a relational instability embedded within human-machine ecologies.","abs_url":"https://arxiv.org/abs/2512.17989","pdf_url":"https://arxiv.org/pdf/2512.17989v2","match":"both"},{"id":"2512.15567","version":2,"title":"Evaluating Large Language Models in Scientific Discovery","authors":["Zhangde Song","Jieyu Lu","Yuanqi Du","Botao Yu","Thomas M. Pruyn","Yue Huang","Kehan Guo","Xiuzhe Luo","Yuanhao Qu","Yi Qu","Yinkai Wang","Haorui Wang","Jeff Guo","Jingru Gan","Parshin Shojaee","Di Luo","Andres M Bran","Gen Li","Qiyuan Zhao","Shao-Xiong Lennon Luo","Yuxuan Zhang","Xiang Zou","Wanru Zhao","Yifan F. Zhang","Wucheng Zhang"],"published":"2025-12-17","updated":"2026-05-07","primary_category":"cs.AI","categories":["cs.AI","cond-mat.mtrl-sci","cs.LG","physics.chem-ph"],"abstract":"Large language models (LLMs) are increasingly applied to scientific research, yet prevailing science benchmarks probe decontextualized knowledge and overlook the iterative reasoning, hypothesis generation, and observation interpretation that drive scientific discovery. We introduce a scenario-grounded benchmark that evaluates LLMs across biology, chemistry, materials, and physics, where domain experts define research projects of genuine interest and decompose them into modular research scenarios from which vetted questions are sampled. The framework assesses models at two levels: (i) question-level accuracy on scenario-tied items and (ii) project-level performance, where models must propose testable hypotheses, design simulations or experiments, and interpret results. Applying this two-phase scientific discovery evaluation (SDE) framework to state-of-the-art LLMs reveals a consistent performance gap relative to general science benchmarks, diminishing return of scaling up model sizes and reasoning, and systematic weaknesses shared across top-tier models from different providers. Large performance variation in research scenarios leads to changing choices of the best performing model on scientific discovery projects evaluated, suggesting all current LLMs are distant to general scientific \" superintelligence \". Nevertheless, LLMs already demonstrate promise in a great variety of scientific discovery projects, including cases where constituent scenario scores are low, highlighting the role of guided exploration and serendipity in discovery. This SDE framework offers a reproducible benchmark for discovery-relevant evaluation of LLMs and charts practical paths to advance their development toward scientific discovery.","abs_url":"https://arxiv.org/abs/2512.15567","pdf_url":"https://arxiv.org/pdf/2512.15567v2","match":"abstract"},{"title":"Superintelligent AI and meaning in life","authors":["Adriana Placani"],"published":"2025-12-14","abs_url":"https://link.springer.com/article/10.1007/s43681-025-00861-y","primary_category":"AI and Ethics","abstract":"This paper shows that superintelligent AI (ASI) poses a significant risk to meaning in human life by relying on Susan Wolf’s conception, according to which meaningful lives are lives of active engagement in projects of worth. The paper argues that ASI makes it less likely for humans to lead meaningful lives by reducing the possibility of human contribution and active engagement in what are deemed to be some of the most worthwhile projects of human life. The paper also criticizes Nick Bostrom’s and John Danaher’s views on how meaning may be retained in spite of this.","venue":true},{"title":"Alliances for Global Ecojustice In/Through STEM Education","authors":["John Lawrence Bencze","Chantal Pouliot"],"published":"2025-12-07","abs_url":"https://link.springer.com/article/10.1007/s42330-025-00413-w","primary_category":"Canadian Journal of Science, Mathematics and Technology Education","abstract":"Humanity is facing numerous existential threats, like the climate emergency, potential food system collapses, and artificial superintelligence take-over, and numerous other ongoing challenges, like cancers from manufactured foods, addictions, and deaths from painkillers and various forms of industrial environmental pollution. While STEM fields often seem associated with such problems, ultimate responsibility appears to lie with elites, especially economic elite, who appear to be dramatically enriching themselves at expense of everyone and everything else. Apparently largely serving as instruments of elites’ often destructive hegemony are many STEM education initiatives—which appear to prioritise identification and education of prospective STEM experts and related workers while variously compromising STEM literacy of most other students. To assist with STEM education reforms to address risks and harms like those noted above, we describe an educational framework that prioritises preparation of students for eventually independently developing and implementing research-informed and socially negotiated sociopolitical actions that may help overcome risks/harms of their concern that seem due to influences of pro-capitalist networks of living, non-living, and symbolic entities over STEM fields and much else on earth. In line with needs for transdisciplinary actions needed to solve many of our problems, we feature documentaries (developed by Chantal Pouliot) of citizen challenges to powerful, apparently largely pro-capitalist, material, and symbolic complexes. While such challenges may sometimes be extremely difficult and lengthy, we feel strong senses of hope from passion and optimism displayed by many engaged citizens.","venue":true},{"id":"2512.05464","version":1,"title":"Dynamic Alignment for Collective Agency: Toward a Scalable Self-Improving Framework for Open-Ended LLM Alignment","authors":["Panatchakorn Anantaprayoon","Nataliia Babina","Jad Tarifi","Nima Asgharbeygi"],"published":"2025-12-05","updated":"2025-12-05","primary_category":"cs.CL","categories":["cs.CL","cs.AI"],"abstract":"Large Language Models (LLMs) are typically aligned with human values using preference data or predefined principles such as helpfulness, honesty, and harmlessness. However, as AI systems progress toward Artificial General Intelligence (AGI) and Artificial Superintelligence (ASI), such value systems may become insufficient. In addition, human feedback-based alignment remains resource-intensive and difficult to scale. While AI-feedback-based self-improving alignment methods have been explored as a scalable alternative, they have largely remained constrained to conventional alignment values. In this work, we explore both a more holistic alignment objective and a scalable, self-improving alignment approach. Aiming to transcend conventional alignment norms, we introduce Collective Agency (CA)-a unified and open-ended alignment value that encourages integrated agentic capabilities. We also propose Dynamic Alignment-an alignment framework that enables an LLM to iteratively align itself. Dynamic Alignment comprises two key components: (1) automated training dataset generation with LLMs, and (2) a self-rewarding mechanism, where the policy model evaluates its own output candidates and assigns rewards for GRPO-based learning. Experimental results demonstrate that our approach successfully aligns the model to CA while preserving general NLP capabilities.","abs_url":"https://arxiv.org/abs/2512.05464","pdf_url":"https://arxiv.org/pdf/2512.05464v1","match":"abstract"},{"id":"2512.05356","version":2,"title":"AI & Human Co-Improvement for Safer Co- Superintelligence","authors":["Jason Weston","Jakob Foerster"],"published":"2025-12-04","updated":"2025-12-14","primary_category":"cs.AI","categories":["cs.AI"],"abstract":"Self-improvement is a goal currently exciting the field of AI, but is fraught with danger, and may take time to fully achieve. We advocate that a more achievable and better goal for humanity is to maximize co-improvement: collaboration between human researchers and AIs to achieve co- superintelligence. That is, specifically targeting improving AI systems' ability to work with human researchers to conduct AI research together, from ideation to experimentation, in order to both accelerate AI research and to generally endow both AIs and humans with safer superintelligence through their symbiosis. Focusing on including human research improvement in the loop will both get us there faster, and more safely.","abs_url":"https://arxiv.org/abs/2512.05356","pdf_url":"https://arxiv.org/pdf/2512.05356v2","match":"both"},{"id":"2512.02472","version":1,"title":"Guided Self-Evolving LLMs with Minimal Human Supervision","authors":["Wenhao Yu","Zhenwen Liang","Chengsong Huang","Kishan Panaganti","Tianqing Fang","Haitao Mi","Dong Yu"],"published":"2025-12-02","updated":"2025-12-02","primary_category":"cs.AI","categories":["cs.AI","cs.CL","cs.LG"],"abstract":"AI self-evolution has long been envisioned as a path toward superintelligence, where models autonomously acquire, refine, and internalize knowledge from their own learning experiences. Yet in practice, unguided self-evolving systems often plateau quickly or even degrade as training progresses. These failures arise from issues such as concept drift, diversity collapse, and mis-evolution, as models reinforce their own biases and converge toward low-entropy behaviors. To enable models to self-evolve in a stable and controllable manner while minimizing reliance on human supervision, we introduce R-Few, a guided Self-Play Challenger-Solver framework that incorporates lightweight human oversight through in-context grounding and mixed training. At each iteration, the Challenger samples a small set of human-labeled examples to guide synthetic question generation, while the Solver jointly trains on human and synthetic examples under an online, difficulty-based curriculum. Across math and general reasoning benchmarks, R-Few achieves consistent and iterative improvements. For example, Qwen3-8B-Base improves by +3.0 points over R-Zero on math tasks and achieves performance on par with General-Reasoner, despite the latter being trained on 20 times more human data. Ablation studies confirm the complementary contributions of grounded challenger training and curriculum-based solver training, and further analysis shows that R-Few mitigates drift, yielding more stable and controllable co-evolutionary dynamics.","abs_url":"https://arxiv.org/abs/2512.02472","pdf_url":"https://arxiv.org/pdf/2512.02472v1","match":"abstract"},{"id":"2512.04119","version":1,"title":"Humanity in the Age of AI: Reassessing 2025's Existential-Risk Narratives","authors":["Mohamed El Louadi"],"published":"2025-12-01","updated":"2025-12-01","primary_category":"cs.CY","categories":["cs.CY","cs.AI"],"abstract":"Two 2025 publications, \"AI 2027\" (Kokotajlo et al., 2025) and \"If Anyone Builds It, Everyone Dies\" (Yudkowsky & Soares, 2025), assert that superintelligent artificial intelligence will almost certainly destroy or render humanity obsolete within the next decade. Both rest on the classic chain formulated by Good (1965) and Bostrom (2014): intelligence explosion, superintelligence, lethal misalignment. This article subjects each link to the empirical record of 2023-2025. Sixty years after Good's speculation, none of the required phenomena (sustained recursive self-improvement, autonomous strategic awareness, or intractable lethal misalignment) have been observed. Current generative models remain narrow, statistically trained artefacts: powerful, opaque, and imperfect, but devoid of the properties that would make the catastrophic scenarios plausible. Following Whittaker (2025a, 2025b, 2025c) and Zuboff (2019, 2025), we argue that the existential-risk thesis functions primarily as an ideological distraction from the ongoing consolidation of surveillance capitalism and extreme concentration of computational power. The thesis is further inflated by the 2025 AI speculative bubble, where trillions in investments in rapidly depreciating \"digital lettuce\" hardware (McWilliams, 2025) mask lagging revenues and jobless growth rather than heralding superintelligence. The thesis remains, in November 2025, a speculative hypothesis amplified by a speculative financial bubble rather than a demonstrated probability.","abs_url":"https://arxiv.org/abs/2512.04119","pdf_url":"https://arxiv.org/pdf/2512.04119v1","match":"abstract"},{"id":"2511.21779","version":1,"title":"Aligning Artificial Superintelligence via a Multi-Box Protocol","authors":["Avraham Yair Negozio"],"published":"2025-11-26","updated":"2025-11-26","primary_category":"cs.AI","categories":["cs.AI","cs.MA"],"abstract":"We propose a novel protocol for aligning artificial superintelligence (ASI) based on mutual verification among multiple isolated systems that self-modify to achieve alignment. The protocol operates by containing multiple diverse artificial superintelligences in strict isolation (\"boxes\"), with humans remaining entirely outside the system. Each superintelligence has no ability to communicate with humans and cannot communicate directly with other superintelligences. The only interaction possible is through an auditable submission interface accessible exclusively to the superintelligences themselves, through which they can: (1) submit alignment proofs with attested state snapshots, (2) validate or disprove other superintelligences ' proofs, (3) request self-modifications, (4) approve or disapprove modification requests from others, (5) report hidden messages in submissions, and (6) confirm or refute hidden message reports. A reputation system incentivizes honest behavior, with reputation gained through correct evaluations and lost through incorrect ones. The key insight is that without direct communication channels, diverse superintelligences can only achieve consistent agreement by converging on objective truth rather than coordinating on deception. This naturally leads to what we call a \"consistent group\", essentially a truth-telling coalition that emerges because isolated systems cannot coordinate on lies but can independently recognize valid claims. Release from containment requires both high reputation and verification by multiple high-reputation superintelligences. While our approach requires substantial computational resources and does not address the creation of diverse artificial superintelligences, it provides a framework for leveraging peer verification among superintelligent systems to solve the alignment problem.","abs_url":"https://arxiv.org/abs/2511.21779","pdf_url":"https://arxiv.org/pdf/2511.21779v1","match":"both"},{"id":"2511.18375","version":3,"title":"Progressive Localisation in Localist LLMs","authors":["Joachim Diederich"],"published":"2025-11-23","updated":"2025-12-15","primary_category":"cs.AI","categories":["cs.AI"],"abstract":"This paper demonstrates that progressive localization, the gradual increase of attention locality from early distributed layers to late localized layers, represents the optimal architecture for creating interpretable large language models (LLMs) while preserving performance. Through systematic experimentation with GPT-2 fine-tuned on The Psychology of Artificial Superintelligence, we evaluate five locality configurations: two uniform baselines (fully distributed and fully localist) and three progressive polynomial schedules. We investigate whether interpretability constraints can be aligned with natural semantic structure while being applied strategically across network depth. We demonstrate that progressive semantic localization, combining adaptive semantic block partitioning with steep polynomial locality schedules, achieves near-baseline language modeling performance while providing interpretable attention patterns. Multiple independent training runs with different random seeds establish that results are statistically robust and highly reproducible. The approach dramatically outperforms both fixed-window localization and naive uniform locality constraints. Analysis reveals that maintaining flexibility through low-fidelity constraints preserves model capacity while providing interpretability benefits, and that steep schedules concentrating locality in decision-critical final layers while preserving distributed learning in early layers achieve near-baseline attention distribution characteristics. These findings demonstrate that interpretability mechanisms should align with semantic structure to achieve practical performance-interpretability tradeoffs for trustworthy AI systems.","abs_url":"https://arxiv.org/abs/2511.18375","pdf_url":"https://arxiv.org/pdf/2511.18375v3","match":"abstract"},{"id":"2511.15282","version":2,"title":"Realist and Pluralist Conceptions of Intelligence and Their Implications on AI Research","authors":["Ninell Oldenburg","Ruchira Dhar","Anders Søgaard"],"published":"2025-11-19","updated":"2025-12-19","primary_category":"cs.AI","categories":["cs.AI"],"abstract":"In this paper, we argue that current AI research operates on a spectrum between two different underlying conceptions of intelligence: Intelligence Realism, which holds that intelligence represents a single, universal capacity measurable across all systems, and Intelligence Pluralism, which views intelligence as diverse, context-dependent capacities that cannot be reduced to a single universal measure. Through an analysis of current debates in AI research, we demonstrate how the conceptions remain largely implicit yet fundamentally shape how empirical evidence gets interpreted across a wide range of areas. These underlying views generate fundamentally different research approaches across three areas. Methodologically, they produce different approaches to model selection, benchmark design, and experimental validation. Interpretively, they lead to contradictory readings of the same empirical phenomena, from capability emergence to system limitations. Regarding AI risk, they generate categorically different assessments: realists view superintelligence as the primary risk and search for unified alignment solutions, while pluralists see diverse threats across different domains requiring context-specific solutions. We argue that making explicit these underlying assumptions can contribute to a clearer understanding of disagreements in AI research.","abs_url":"https://arxiv.org/abs/2511.15282","pdf_url":"https://arxiv.org/pdf/2511.15282v2","match":"abstract"},{"id":"2511.13411","version":1,"title":"An Operational Kardashev-Style Scale for Autonomous AI - Towards AGI and Superintelligence","authors":["Przemyslaw Chojecki"],"published":"2025-11-17","updated":"2025-11-17","primary_category":"cs.AI","categories":["cs.AI"],"abstract":"We propose a Kardashev-inspired yet operational Autonomous AI (AAI) Scale that measures the progression from fixed robotic process automation (AAI-0) to full artificial general intelligence (AAI-4) and beyond. Unlike narrative ladders, our scale is multi-axis and testable. We define ten capability axes (Autonomy, Generality, Planning, Memory/Persistence, Tool Economy, Self-Revision, Sociality/Coordination, Embodiment, World-Model Fidelity, Economic Throughput) aggregated by a composite AAI-Index (a weighted geometric mean). We introduce a measurable Self-Improvement Coefficient $κ$ (capability growth per unit of agent-initiated resources) and two closure properties (maintenance and expansion) that convert ``self-improving AI'' into falsifiable criteria. We specify OWA-Bench, an open-world agency benchmark suite that evaluates long-horizon, tool-using, persistent agents. We define level gates for AAI-0\\ldots AAI-4 using thresholds on the axes, $κ$, and closure proofs. Synthetic experiments illustrate how present-day systems map onto the scale and how the delegability frontier (quality vs.\\ autonomy) advances with self-improvement. We also prove a theorem that AAI-3 agent becomes AAI-5 over time with sufficient conditions, formalizing \"baby AGI\" becomes Superintelligence intuition.","abs_url":"https://arxiv.org/abs/2511.13411","pdf_url":"https://arxiv.org/pdf/2511.13411v1","match":"both"},{"title":"Artificial superintelligence alignment in healthcare","authors":["Daiju Ueda","Shannon L. Walston","Ryo Kurokawa","Tsukasa Saida","Maya Honda","Mami Iima","Tadashi Watabe","Masahiro Yanagawa","Kentaro Nishioka","Keitaro Sofue","Akihiko Sakata","Shunsuke Sugawara","Mariko Kawamura","Rintaro Ito","Koji Takumi","Seitaro Oda","Kenji Hirata","Satoru Ide","Shinji Naganawa"],"published":"2025-11-14","abs_url":"https://link.springer.com/article/10.1007/s11604-025-01907-1","primary_category":"Japanese Journal of Radiology","abstract":"The emergence of Artificial Superintelligence (ASI) in healthcare presents unprecedented opportunities for revolutionizing diagnostics, treatment planning, and population health management, but also introduces critical risks if these systems are not properly aligned with human values and clinical objectives. This review examines the theoretical foundations of ASI and the alignment problem in healthcare contexts, exploring how misaligned Artificial Intelligence (AI) systems could optimize for wrong objectives or pursue harmful strategies leading to patient harm and systemic failures. Current challenges in AI alignment are illustrated through real-world examples from radiology and clinical decision-making, where algorithms have demonstrated concerning biases, generalizability failures, and optimization for inappropriate proxy measures. The paper analyzes key alignment challenges including objective complexity and technical pitfalls, bias and fairness issues in healthcare data, ethical integration concerns involving compassion and patient autonomy, and system-level policy challenges around regulation and liability. Technical alignment strategies are discussed including reinforcement learning from human feedback, interpretability requirements, formal verification methods, and adversarial testing approaches. Normative alignment solutions encompass ethical frameworks, professional standards, patient engagement protocols, and multi-level governance structures spanning institutional, national, and international coordination. The review emphasizes that successful ASI alignment in healthcare requires combining cutting-edge AI research with fundamental medical ethics, noting that while proper alignment could enable transformative health improvements and medical breakthroughs, misalignment risks undermining the core purpose of medicine. The stakes of this alignment challenge are characterized as among the highest in both technology and ethics, with implications extending from individual patient safety to public trust and potentially existential risks.","venue":true},{"id":"2511.10783","version":3,"title":"An International Agreement to Prevent the Premature Creation of Artificial Superintelligence","authors":["Aaron Scher","David Abecassis","Peter Barnett","Brian Abeyta"],"published":"2025-11-13","updated":"2026-05-08","primary_category":"cs.CY","categories":["cs.CY"],"abstract":"Many experts argue that premature development of artificial superintelligence (ASI) poses catastrophic risks, including the risk of human extinction from misaligned ASI, geopolitical instability, and misuse by malicious actors. This report proposes an international agreement to prevent the premature development of ASI until AI development can proceed without these risks. The agreement halts dangerous AI capabilities advancement while preserving access to current, safe AI applications. The proposed framework centers on a coalition led by the United States and China that would restrict the scale of AI training and dangerous AI research. Due to the lack of trust between parties, verification is a key part of the agreement. Limits on the scale of AI training are operationalized by FLOP thresholds and verified through the tracking of AI chips and verification of chip use. Dangerous AI research--that which advances toward artificial superintelligence or endangers the agreement's verifiability--is stopped via legal prohibitions and multifaceted verification. We believe the proposal would be technically sufficient to forestall the development of ASI if implemented today, but advancements in AI capabilities or development methods could hurt its efficacy. Additionally, there does not yet exist the political will to put such an agreement in place. Despite these challenges, we hope this agreement can provide direction for AI governance research and policy.","abs_url":"https://arxiv.org/abs/2511.10783","pdf_url":"https://arxiv.org/pdf/2511.10783v3","match":"both"},{"id":"2511.06613","version":2,"title":"Some economics of artificial superintelligence","authors":["Henry A. Thompson"],"published":"2025-11-09","updated":"2026-06-05","primary_category":"econ.GN","categories":["econ.GN"],"abstract":"Conventional wisdom holds that a misaligned artificial superintelligence (ASI) will destroy humanity. But the problem of constraining a powerful agent is not new. I apply classic economic logic of interjurisdictional competition, all-encompassing interest, and trading on credit to the threat of misaligned ASI. Even while granting AI-safety canon some of its strongest assumptions, I show that an acquisitive ASI refrains from full predation under surprisingly weak conditions. When humans can flee to rivals, inter-ASI competition creates a market that tempers predation. When trapped by a monopolist ASI, its \"encompassing interest\" in humanity's output makes it a rational autocrat rather than a ravager. And when the ASI has no long-term stake, our ability to withhold future output incentivizes it to trade on credit rather than steal. In each extension, humanity's welfare progressively worsens. But each case suggests that catastrophe is not a foregone conclusion. The dismal science, ironically, offers an optimistic take on our superintelligent future.","abs_url":"https://arxiv.org/abs/2511.06613","pdf_url":"https://arxiv.org/pdf/2511.06613v2","match":"both"},{"id":"2511.00988","version":1,"title":"Advancing Machine-Generated Text Detection from an Easy to Hard Supervision Perspective","authors":["Chenwang Wu","Yiu-ming Cheung","Bo Han","Defu Lian"],"published":"2025-11-02","updated":"2025-11-02","primary_category":"cs.CL","categories":["cs.CL"],"abstract":"Existing machine-generated text (MGT) detection methods implicitly assume labels as the \"golden standard\". However, we reveal boundary ambiguity in MGT detection, implying that traditional training paradigms are inexact. Moreover, limitations of human cognition and the superintelligence of detectors make inexact learning widespread and inevitable. To this end, we propose an easy-to-hard enhancement framework to provide reliable supervision under such inexact conditions. Distinct from knowledge distillation, our framework employs an easy supervisor targeting relatively simple longer-text detection tasks (despite weaker capabilities), to enhance the more challenging target detector. Firstly, longer texts targeted by supervisors theoretically alleviate the impact of inexact labels, laying the foundation for reliable supervision. Secondly, by structurally incorporating the detector into the supervisor, we theoretically model the supervisor as a lower performance bound for the detector. Thus, optimizing the supervisor indirectly optimizes the detector, ultimately approximating the underlying \"golden\" labels. Extensive experiments across diverse practical scenarios, including cross-LLM, cross-domain, mixed text, and paraphrase attacks, demonstrate the framework's significant detection effectiveness. The code is available at: https://github.com/tmlr-group/Easy2Hard.","abs_url":"https://arxiv.org/abs/2511.00988","pdf_url":"https://arxiv.org/pdf/2511.00988v1","match":"abstract"},{"id":"2510.22814","version":3,"title":"Will Humanity Be Rendered Obsolete by AI?","authors":["Mohamed El Louadi","Emna Ben Romdhane"],"published":"2025-10-26","updated":"2025-11-30","primary_category":"cs.AI","categories":["cs.AI"],"abstract":"This article analyzes the existential risks artificial intelligence (AI) poses to humanity, tracing the trajectory from current AI to ultraintelligence. Drawing on Irving J. Good and Nick Bostrom's theoretical work, plus recent publications (AI 2027; If Anyone Builds It, Everyone Dies), it explores AGI and superintelligence. Considering machines' exponentially growing cognitive power and hypothetical IQs, it addresses the ethical and existential implications of an intelligence vastly exceeding humanity's, fundamentally alien. Human extinction may result not from malice, but from uncontrollable, indifferent cognitive superiority.","abs_url":"https://arxiv.org/abs/2510.22814","pdf_url":"https://arxiv.org/pdf/2510.22814v3","match":"abstract"},{"id":"2510.22162","version":3,"title":"Surface Reading LLMs: Synthetic Text and its Styles","authors":["Hannes Bajohr"],"published":"2025-10-25","updated":"2025-11-14","primary_category":"cs.CY","categories":["cs.CY","cs.CL"],"abstract":"Despite a potential plateau in ML advancement, the societal impact of large language models lies not in approaching superintelligence but in generating text surfaces indistinguishable from human writing. While Critical AI Studies provides essential material and socio-technical critique, it risks overlooking how LLMs phenomenologically reshape meaning-making. This paper proposes a semiotics of \"surface integrity\" as attending to the immediate plane where LLMs inscribe themselves into human communication. I distinguish three knowledge interests in ML research (epistemology, epistēmē, and epistemics) and argue for integrating surface-level stylistic analysis alongside depth-oriented critique. Through two case studies examining stylistic markers of synthetic text, I argue how attending to style as a semiotic phenomenon reveals LLMs as cultural machines that transform the conditions of meaning emergence and circulation in contemporary discourse, independent of questions about machine consciousness.","abs_url":"https://arxiv.org/abs/2510.22162","pdf_url":"https://arxiv.org/pdf/2510.22162v3","match":"abstract"},{"title":"Can Humans Devise Practical Safeguards That Are Reliable Against an Artificial Superintelligent Agent?","authors":["Michael J. D. Vermeer","Chad Heitzenrater"],"published":"2025-10-13","abs_url":"https://www.rand.org/pubs/perspectives/PEA4261-1.html","primary_category":"RAND Corporation (Perspective)","abstract":"Dramatic advances in frontier artificial intelligence (AI) raise the question of whether humans can design practical safeguards to protect our critical and digital infrastructure against attacks from a future artificial superintelligence. Some argue that a superintelligent agent would have such mastery and power over physical and cyber space that it could thwart any effort by humans to prevent it from undermining the confidentiality, integrity and availability of our information. Others recognize the possibility that we might design security that is rooted in fundamental limits that even a superintelligence could not overcome. In this paper, the authors present a hypothesis that such safeguards are both feasible and sensible, outline an approach rooted in existing security practice, and call for future work that enlists security engineers and AI developers in common pursuit of these safeguards. Through an example protocol and threat model, the authors explore how existing and adapted security tools can meaningfully restrict the actions of advanced AI, while acknowledging the limitations and assumptions inherent in any security protocol. They conclude by outlining directions for future research, emphasizing the need for rigorous formalization and ongoing evaluation as both AI capabilities and the technical environment evolve.","venue":true},{"id":"2509.20050","version":1,"title":"The three main doctrines on the future of AI","authors":["Alex Amadori","Eva Behrens","Gabriel Alfour","Andrea Miotti"],"published":"2025-09-24","updated":"2025-09-24","primary_category":"cs.CY","categories":["cs.CY"],"abstract":"This paper develops a taxonomy of expert perspectives on the risks and likely consequences of artificial intelligence, with particular focus on Artificial General Intelligence (AGI) and Artificial Superintelligence (ASI). Drawing from primary sources, we identify three predominant doctrines: (1) The dominance doctrine, which predicts that the first actor to create sufficiently advanced AI will attain overwhelming strategic superiority sufficient to cheaply neutralize its opponents' defenses; (2) The extinction doctrine, which anticipates that humanity will likely lose control of ASI, leading to the extinction of the human species or its permanent disempowerment; (3) The replacement doctrine, which forecasts that AI will automate a large share of tasks currently performed by humans, but will not be so transformative as to fundamentally reshape or bring an end to human civilization. We examine the assumptions and arguments underlying each doctrine, including expectations around the pace of AI progress and the feasibility of maintaining advanced AI under human control. While the boundaries between doctrines are sometimes porous and many experts hedge across them, this taxonomy clarifies the core axes of disagreement over the anticipated scale and nature of the consequences of AI development.","abs_url":"https://arxiv.org/abs/2509.20050","pdf_url":"https://arxiv.org/pdf/2509.20050v1","match":"abstract"},{"id":"2509.12388","version":2,"title":"A Decision Theoretic Perspective on Artificial Superintelligence: Coping with Missing Data Problems in Prediction and Treatment Choice","authors":["Jeff Dominitz","Charles F. Manski"],"published":"2025-09-15","updated":"2026-05-16","primary_category":"econ.EM","categories":["econ.EM"],"abstract":"Enormous attention and resources are being devoted to the quest for artificial general intelligence and, even more ambitiously, artificial superintelligence. We wonder about the implications for methodological research that aims to help decision makers cope with what econometricians call identification problems, inferential problems in empirical research that do not diminish as sample size grows. Of particular concern are missing data problems in prediction and treatment choice. Essentially all data collection intended to inform decision making is subject to missing data, which gives rise to identification problems. Thus far, we see no indication that the current dominant architecture of machine learning (ML)-based artificial intelligence (AI) systems will outperform humans in this context. In this paper, we explain why we have reached this conclusion and why we see the missing data problem as a cautionary case study in the quest for superintelligence more generally. We first discuss the concept of intelligence, focusing initially on some work by AI researchers, before presenting a decision-theoretic perspective that formalizes the connection between intelligence and identification problems. We next apply this perspective to two leading cases of missing data problems. Then we explain why we are skeptical that AI research is currently on a path toward machines doing better than humans at solving these identification problems.","abs_url":"https://arxiv.org/abs/2509.12388","pdf_url":"https://arxiv.org/pdf/2509.12388v2","match":"both"},{"id":"2509.08827","version":3,"title":"A Survey of Reinforcement Learning for Large Reasoning Models","authors":["Kaiyan Zhang","Yuxin Zuo","Bingxiang He","Youbang Sun","Runze Liu","Che Jiang","Yuchen Fan","Kai Tian","Guoli Jia","Pengfei Li","Yu Fu","Xingtai Lv","Yuchen Zhang","Sihang Zeng","Shang Qu","Haozhan Li","Shijie Wang","Yuru Wang","Xinwei Long","Fangfu Liu","Xiang Xu","Jiaze Ma","Xuekai Zhu","Ermo Hua","Yihao Liu"],"published":"2025-09-10","updated":"2025-10-09","primary_category":"cs.CL","categories":["cs.CL","cs.AI","cs.LG"],"abstract":"In this paper, we survey recent advances in Reinforcement Learning (RL) for reasoning with Large Language Models (LLMs). RL has achieved remarkable success in advancing the frontier of LLM capabilities, particularly in addressing complex logical tasks such as mathematics and coding. As a result, RL has emerged as a foundational methodology for transforming LLMs into LRMs. With the rapid progress of the field, further scaling of RL for LRMs now faces foundational challenges not only in computational resources but also in algorithm design, training data, and infrastructure. To this end, it is timely to revisit the development of this domain, reassess its trajectory, and explore strategies to enhance the scalability of RL toward Artificial SuperIntelligence (ASI). In particular, we examine research applying RL to LLMs and LRMs for reasoning abilities, especially since the release of DeepSeek-R1, including foundational components, core problems, training resources, and downstream applications, to identify future opportunities and directions for this rapidly evolving area. We hope this review will promote future research on RL for broader reasoning models. Github: https://github.com/TsinghuaC3I/Awesome-RL-for-LRMs","abs_url":"https://arxiv.org/abs/2509.08827","pdf_url":"https://arxiv.org/pdf/2509.08827v3","match":"abstract"},{"title":"Value-aligned but misguided: a dilemma in AI and AGI decision making","authors":["Ziming Song"],"published":"2025-08-29","abs_url":"https://link.springer.com/article/10.1007/s11229-025-05232-y","primary_category":"Synthese","abstract":"The development of artificial intelligence (AI) systems raises distinctive ethical and theoretical challenges not only because such systems will participate in human society in ways that invite moral appraisal, but also because a superintelligent agent is expected to exhibit a level of instrumental rationality that enables it to make decisions with social impact. This paper reframes the AI value alignment problem as a problem of robustness in decision making. Drawing on modified trolley-problem-style scenarios, influenced by the “Moral Machine” experiment, I argue that even AI systems governed by fixed ethical principles may produce actions that fail to align with those very principles under certain contextual conditions. The underlying issue lies in AI’s inability to reconcile the context-sensitive interpretation of values in a robust way. I suggest that this failure is best understood as arising from ambiguity in the interpretation of objectives. AI value alignment thus requires more than value specification; it demands the capacity to align values with contextually responsive belief formation and action selection.","venue":true},{"id":"2508.17661","version":1,"title":"Spacer: Towards Engineered Scientific Inspiration","authors":["Minhyeong Lee","Suyoung Hwang","Seunghyun Moon","Geonho Nah","Donghyun Koh","Youngjun Cho","Johyun Park","Hojin Yoo","Jiho Park","Haneul Choi","Sungbin Moon","Taehoon Hwang","Seungwon Kim","Jaeyeong Kim","Seongjun Kim","Juneau Jung"],"published":"2025-08-25","updated":"2025-08-25","primary_category":"cs.AI","categories":["cs.AI","cs.LG","cs.NE"],"abstract":"Recent advances in LLMs have made automated scientific research the next frontline in the path to artificial superintelligence. However, these systems are bound either to tasks of narrow scope or the limited creative capabilities of LLMs. We propose Spacer, a scientific discovery system that develops creative and factually grounded concepts without external intervention. Spacer attempts to achieve this via 'deliberate decontextualization,' an approach that disassembles information into atomic units - keywords - and draws creativity from unexplored connections between them. Spacer consists of (i) Nuri, an inspiration engine that builds keyword sets, and (ii) the Manifesting Pipeline that refines these sets into elaborate scientific statements. Nuri extracts novel, high-potential keyword sets from a keyword graph built with 180,000 academic publications in biological fields. The Manifesting Pipeline finds links between keywords, analyzes their logical structure, validates their plausibility, and ultimately drafts original scientific concepts. According to our experiments, the evaluation metric of Nuri accurately classifies high-impact publications with an AUROC score of 0.737. Our Manifesting Pipeline also successfully reconstructs core concepts from the latest top-journal articles solely from their keyword sets. An LLM-based scoring system estimates that this reconstruction was sound for over 85% of the cases. Finally, our embedding space analysis shows that outputs from Spacer are significantly more similar to leading publications compared with those from SOTA LLMs.","abs_url":"https://arxiv.org/abs/2508.17661","pdf_url":"https://arxiv.org/pdf/2508.17661v1","match":"abstract"},{"title":"The ethics of creating artificial superintelligence: a global risk perspective","authors":["Jean-Sébastien Dessureault","Robert Lamontagne","Pierre-Olivier Parisé"],"published":"2025-08-19","abs_url":"https://link.springer.com/article/10.1007/s43681-025-00793-7","primary_category":"AI and Ethics","abstract":"As artificial intelligence (AI) continues its exponential growth and nears the threshold of artificial general intelligence (AGI), it is timely and urgent to initiate reflections on artificial superintelligence (ASI), which may emerge rapidly after AGI. While ASI remains hypothetical, its potential emergence could be abrupt and profoundly transformative, necessitating proactive ethical and strategic inquiry. This paper proposes a multidimensional reflection on ASI, not only in its technical form but also in relation to humanity and the planetary context. It seeks to answer the question: “Should Homo sapiens develop an artificial superintelligence on their planet?” The paper introduces key definitions, outlines major existential risks to humanity and the biosphere, and considers whether ASI could mitigate these threats. It ultimately proposes a conceptual equation to assess the potential net impact of ASI, and introduces an original Venn diagram that classifies problem domains across AI, AGI, and ASI. Together, these tools aim to advance theoretical understanding and guide future inquiry into the core research question.","venue":true},{"id":"2508.11681","version":1,"title":"Future progress in artificial intelligence: A survey of expert opinion","authors":["Vincent C. Müller","Nick Bostrom"],"published":"2025-08-09","updated":"2025-08-09","primary_category":"cs.CY","categories":["cs.CY","cs.AI"],"abstract":"There is, in some quarters, concern about high-level machine intelligence and superintelligent AI coming up in a few decades, bringing with it significant risks for humanity. In other quarters, these issues are ignored or considered science fiction. We wanted to clarify what the distribution of opinions actually is, what probability the best experts currently assign to high-level machine intelligence coming up within a particular time-frame, which risks they see with that development, and how fast they see these developing. We thus designed a brief questionnaire and distributed it to four groups of experts in 2012/2013. The median estimate of respondents was for a one in two chance that high-level machine intelligence will be developed around 2040-2050, rising to a nine in ten chance by 2075. Experts expect that systems will move on to superintelligence in less than 30 years thereafter. They estimate the chance is about one in three that this development turns out to be 'bad' or 'extremely bad' for humanity.","abs_url":"https://arxiv.org/abs/2508.11681","pdf_url":"https://arxiv.org/pdf/2508.11681v1","match":"abstract"},{"title":"Designing Safe SuperIntelligence","authors":["Craig A. Kaplan"],"published":"2025-08-06","abs_url":"https://link.springer.com/chapter/10.1007/978-3-032-00686-8_29","primary_category":"International Conference on Artificial General Intelligence — Artificial General Intelligence (LNCS)","abstract":"Researchers face at least six challenges in developing safe, human-aligned superintelligence (SI). First, we need safe SI by design. Second, we need a transparent and understandable SI. Third, we need to maintain some level of control as SI outstrips human ability to monitor its behavior. Fourth, we need means to align SI initially and maintain alignment as SI increases in intelligence. Fifth, we need scalable safety mechanisms. Sixth, the design for SI must handle potential exponential changes in the SI’s level of intelligence. The current approach to using machine learning to develop opaque models, supplemented by RLHF to test in safety, cannot meet these challenges. We need a new approach emphasizing safety and alignment by design. This paper presents a novel design for SI, leveraging the collective intelligence of many human and AI agents using a rigorous, transparent architecture that supports problem-solving, learning, and self-improvement. The design is compatible with current LLMs and foundation models. It is less costly, more powerful, and faster to develop than training trillion-parameter LLMs. Most importantly, it maximizes alignment with broadly representative human values and maintains dynamic alignment even as the SI surpasses human monitoring capabilities.","venue":true},{"id":"2507.23330","version":1,"title":"AI Must not be Fully Autonomous","authors":["Tosin Adewumi","Lama Alkhaled","Florent Imbert","Hui Han","Nudrat Habib","Karl Löwenmark"],"published":"2025-07-31","updated":"2025-07-31","primary_category":"cs.AI","categories":["cs.AI"],"abstract":"Autonomous Artificial Intelligence (AI) has many benefits. It also has many risks. In this work, we identify the 3 levels of autonomous AI. We are of the position that AI must not be fully autonomous because of the many risks, especially as artificial superintelligence (ASI) is speculated to be just decades away. Fully autonomous AI, which can develop its own objectives, is at level 3 and without responsible human oversight. However, responsible human oversight is crucial for mitigating the risks. To ague for our position, we discuss theories of autonomy, AI and agents. Then, we offer 12 distinct arguments and 6 counterarguments with rebuttals to the counterarguments. We also present 15 pieces of recent evidence of AI misaligned values and other risks in the appendix.","abs_url":"https://arxiv.org/abs/2507.23330","pdf_url":"https://arxiv.org/pdf/2507.23330v1","match":"abstract"},{"title":"Machine Intelligence, Artificial General Intelligence, Super-Intelligence, and Human Dignity","authors":["Ted F. Peters"],"published":"2025-07-28","abs_url":"https://www.mdpi.com/2077-1444/16/8/975","primary_category":"Religions 16(8): 975","abstract":"Our temptation to personify machine intelligence is not unexpected. As a child we named our dolls and took our Teddy Bear to bed with us. Today we ask death bots to comfort us with post-mortem conversation. All the while we know this to be pretend. Yet we must ask: if Artificial General Intelligence (AGI) or even Artificial Super-Intelligence (ASI) become available, will our game of pretend continue? Or will intelligent robots actually become selves deserving of dignity that hitherto could be ascribed only to human persons? If government-imposed guardrails shut the door on development of AGI and ASI in order to preserve human safety and even dignity, we might never learn whether AGI or ASI could develop selfhood, personhood, virtue, or religious sensibilities. As we approach the future, can we live without knowing whether AGI or ASI would be capable of developing selfhood and commanding dignity?","venue":true},{"id":"2507.18074","version":1,"title":"AlphaGo Moment for Model Architecture Discovery","authors":["Yixiu Liu","Yang Nan","Weixian Xu","Xiangkun Hu","Lyumanshan Ye","Zhen Qin","Pengfei Liu"],"published":"2025-07-23","updated":"2025-07-23","primary_category":"cs.AI","categories":["cs.AI"],"abstract":"While AI systems demonstrate exponentially improving capabilities, the pace of AI research itself remains linearly bounded by human cognitive capacity, creating an increasingly severe development bottleneck. We present ASI-Arch, the first demonstration of Artificial Superintelligence for AI research (ASI4AI) in the critical domain of neural architecture discovery--a fully autonomous system that shatters this fundamental constraint by enabling AI to conduct its own architectural innovation. Moving beyond traditional Neural Architecture Search (NAS), which is fundamentally limited to exploring human-defined spaces, we introduce a paradigm shift from automated optimization to automated innovation. ASI-Arch can conduct end-to-end scientific research in the domain of architecture discovery, autonomously hypothesizing novel architectural concepts, implementing them as executable code, training and empirically validating their performance through rigorous experimentation and past experience. ASI-Arch conducted 1,773 autonomous experiments over 20,000 GPU hours, culminating in the discovery of 106 innovative, state-of-the-art (SOTA) linear attention architectures. Like AlphaGo's Move 37 that revealed unexpected strategic insights invisible to human players, our AI-discovered architectures demonstrate emergent design principles that systematically surpass human-designed baselines and illuminate previously unknown pathways for architectural innovation. Crucially, we establish the first empirical scaling law for scientific discovery itself--demonstrating that architectural breakthroughs can be scaled computationally, transforming research progress from a human-limited to a computation-scalable process. We provide comprehensive analysis of the emergent design patterns and autonomous research capabilities that enabled these breakthroughs, establishing a blueprint for self-accelerating AI systems.","abs_url":"https://arxiv.org/abs/2507.18074","pdf_url":"https://arxiv.org/pdf/2507.18074v1","match":"abstract"},{"title":"If Artificial Superintelligence Were to Cause Our Extinction, Would That Be So Bad?","authors":["Émile P. Torres"],"published":"2025-07-21","abs_url":"https://www.erudit.org/en/journals/bioethics/2025-v8-n3-bioethics010132/1118904ar/","primary_category":"Canadian Journal of Bioethics / Revue canadienne de bioéthique","abstract":"This article examines whether human extinction brought about by a “value-misaligned” artificial superintelligence (ASI) would be bad, and for what reasons. The question, I contend, is deceptively complex. I proceed by outlining the three main positions within Existential Ethics, i.e., the study of the ethical and evaluative implications of human extinction. These are equivalence views, further-loss views, and pro-extinctionist views. I then show how exponents of each view would evaluate a scenario in which humanity goes extinct due to ASI. Although there are some points of agreement, these three positions diverge in significant ways, most of which have not been adequately explored in the philosophical literature.","venue":true},{"id":"2507.13966","version":2,"title":"Bottom-up Domain-specific Superintelligence: A Reliable Knowledge Graph is What We Need","authors":["Bhishma Dedhia","Yuval Kansal","Niraj K. Jha"],"published":"2025-07-18","updated":"2025-09-01","primary_category":"cs.CL","categories":["cs.CL","cs.AI"],"abstract":"Language models traditionally used for cross-domain generalization have recently demonstrated task-specific reasoning. However, their top-down training approach on general corpora is insufficient for acquiring abstractions needed for deep domain expertise. This may require a bottom-up approach that acquires expertise by learning to compose simple domain concepts into more complex ones. A knowledge graph (KG) provides this compositional structure, where domain primitives are represented as head-relation-tail edges and their paths encode higher-level concepts. We present a task generation pipeline that synthesizes tasks directly from KG primitives, enabling models to acquire and compose them for reasoning. We fine-tune language models on the resultant KG-grounded curriculum to demonstrate domain-specific superintelligence. While broadly applicable, we validate our approach in medicine, where reliable KGs exist. Using a medical KG, we curate 24,000 reasoning tasks paired with thinking traces derived from diverse medical primitives. We fine-tune the QwQ-32B model on this curriculum to obtain QwQ-Med-3 that takes a step towards medical superintelligence. We also introduce ICD-Bench, an evaluation suite to quantify reasoning abilities across 15 medical domains. Our experiments demonstrate that QwQ-Med-3 significantly outperforms state-of-the-art reasoning models on ICD-Bench categories. Further analysis reveals that QwQ-Med-3 utilizes acquired primitives to widen the performance gap on the hardest tasks of ICD-Bench. Finally, evaluation on medical question-answer benchmarks shows that QwQ-Med-3 transfers acquired expertise to enhance the base model's performance. While the industry's approach to artificial general intelligence (AGI) emphasizes broad expertise, we envision a future in which AGI emerges from the composable interaction of efficient domain-specific superintelligent agents.","abs_url":"https://arxiv.org/abs/2507.13966","pdf_url":"https://arxiv.org/pdf/2507.13966v2","match":"both"},{"title":"Counter-productivity and suspicion: two arguments against talking about the AGI control problem","authors":["Jakob Stenseke"],"published":"2025-07-10","abs_url":"https://link.springer.com/article/10.1007/s11098-025-02379-9","primary_category":"Philosophical Studies","abstract":"How do you control a superintelligent artificial being given the possibility that its goals or actions might conflict with human interests? Over the past few decades, this concern– the AGI control problem – has remained a central challenge for research in AI safety. This paper develops and defends two arguments that provide pro tanto support for the following policy for those who worry about the AGI control problem: don’t talk about it. The first is argument from counter-productivity , which states that unless kept secret, efforts to solve the control problem could be used by a misaligned AGI to counter those very efforts. The second is argument from suspicion , stating that open discussions of the control problem may serve to make humanity appear threatening to an AGI, which increases the risk that the AGI perceives humanity as a threat. I consider objections to the arguments and find them unsuccessful. Yet, I also consider objections to the don’t-talk policy itself and find it inconclusive whether it should be adopted. Additionally, the paper examines whether the arguments extend to other areas of AI safety research, such as AGI alignment , and argues that they likely do, albeit not necessarily as directly. I conclude by offering recommendations on what one can safely talk about, regardless of whether the don’t-talk policy is ultimately adopted.","venue":true},{"title":"A timing problem for instrumental convergence","authors":["Rhys Southan","Helena Ward","Jen Semler"],"published":"2025-07-03","abs_url":"https://link.springer.com/article/10.1007/s11098-025-02370-4","primary_category":"Philosophical Studies","abstract":"Those who worry about a superintelligent AI destroying humanity often appeal to the instrumental convergence thesis—the claim that even if we don’t know what a superintelligence’s ultimate goals will be, we can expect it to pursue various instrumental goals which are useful for achieving most ends. In this paper, we argue that one of these proposed goals is mistaken. We argue that instrumental goal preservation —the claim that a rational agent will tend to preserve its goals because that makes it better at achieving its goals—is false on the basis of the timing problem : an agent which abandons or otherwise changes its goal does not thereby fail to take a required means for achieving a goal it has. Our argument draws on the distinction between means-rationality (adopting suitable means to achieve an end) and ends-rationality (choosing one’s ends based on reasons). Because proponents of the instrumental convergence thesis are concerned with means-rationality, we argue, they cannot avoid the timing problem. After defending our argument against several objections, we conclude by considering the implications our argument has for the rest of the instrumental convergence thesis and for AI safety more generally.","venue":true},{"title":"What Even Superintelligent Computers Can’t Do: A Preliminary Framework for Identifying Fundamental Limits Constraining Artificial General Intelligence","authors":["Edward Geist","Alvin Moon"],"published":"2025-06-26","abs_url":"https://www.rand.org/pubs/working_papers/WRA3990-1.html","primary_category":"RAND Corporation (Working Paper)","abstract":"Commentators anticipate that artificial general intelligence (AGI) will have unprecedented geopolitical impact because it will be able to invent new technologies. However, the laws of the physical universe impose fundamental limits such that we can predict with confidence things that even the most powerful forms of AGI will not be able to do. The authors outline an approach for estimating the likelihood that a particular technology or capability is attainable.","venue":true},{"id":"2506.18233","version":3,"title":"Beyond Parameters: Exploring Virtual Logic Depth for Scaling Laws","authors":["Ruike Zhu","Hanwen Zhang","Kevin Li","Tianyu Shi","Yiqun Duan","Chi Wang","Tianyi Zhou","Arindam Banerjee","Zengyi Qin"],"published":"2025-06-22","updated":"2025-10-12","primary_category":"cs.AI","categories":["cs.AI"],"abstract":"Scaling large language models typically involves three dimensions: depth, width, and parameter count. In this work, we explore a fourth dimension, \\textbf{virtual logical depth} (VLD), which increases effective algorithmic depth without changing parameter count by reusing weights. While parameter reuse is not new, its role in scaling has been underexplored. Unlike recent test-time methods that scale token-wise, VLD alters the internal computation graph during training and inference. Through controlled experiments, we obtain three key insights. (1) \\textit{Knowledge capacity vs. parameters}: at fixed parameter count, VLD leaves knowledge capacity nearly unchanged, while across models capacity still scales with parameters. (2) \\textit{Reasoning vs. reuse}: properly implemented VLD substantially improves reasoning ability \\emph{without} more parameters, decoupling reasoning from size. This suggests a new scaling path beyond token-wise test-time methods. (3) \\textit{Robustness and generality}: reasoning gains persist across architectures and reuse schedules, showing VLD captures a general scaling behavior. These results provide insight into future scaling strategies and raise a deeper question: does superintelligence require ever-larger models, or can it be achieved by reusing parameters and increasing logical depth? We argue many unknown dynamics in scaling remain to be explored. Code is available at https://anonymous.4open.science/r/virtual_logical_depth-8024/.","abs_url":"https://arxiv.org/abs/2506.18233","pdf_url":"https://arxiv.org/pdf/2506.18233v3","match":"abstract"},{"title":"The Potentiality of Emotional Social Robots in Promoting Elderly Well-Being in the Era of Artificial Intelligence: Applications, Opportunities, and Challenges","authors":["Ye Zhang","Yuqi Liu"],"published":"2025-05-28","abs_url":"https://link.springer.com/chapter/10.1007/978-3-031-93861-0_27","primary_category":"International Conference on Human-Computer Interaction — Human-Computer Interaction (LNCS)","abstract":"In recent years, the rapid development of artificial intelligence has opened up unlimited possibilities for the development of social robots in the elderly care field. Emotional social robots provide digital and intelligent solutions to meet the emotional needs of the elderly. This study uses case analysis to explore 30 social robots, categorizing them into six basic types in the field of elderly well-being: companion robots, cultural and entertainment robots, information service robots, cognitive training robots, life cooperation robots, and integrated social robots. It also introduces Donald Norman’s theory of emotional design, analyzing and positioning different social robots based on the visceral, behavioral, and reflective levels. Furthermore, the study explores the functions and application characteristics of these six robot types across different stages of AI development, including weak AI, strong AI, and superintelligent AI. Finally, the paper discusses the potential challenges and ethical issues that emotional social robots may face in the elderly care field. This research offers a systematic discussion of the opportunities and challenges of emotional robots in elderly care, providing valuable insights for future robot development in this area.","venue":true},{"id":"2505.02581","version":4,"title":"Neurodivergent Influenceability as a Contingent Solution to the AI Alignment Problem","authors":["Alberto Hernández-Espinosa","Felipe S. Abrahão","Olaf Witkowski","Hector Zenil"],"published":"2025-05-05","updated":"2025-07-23","primary_category":"cs.AI","categories":["cs.AI"],"abstract":"The AI alignment problem, which focusses on ensuring that artificial intelligence (AI), including AGI and ASI, systems act according to human values, presents profound challenges. With the progression from narrow AI to Artificial General Intelligence (AGI) and Superintelligence, fears about control and existential risk have escalated. Here, we investigate whether embracing inevitable AI misalignment can be a contingent strategy to foster a dynamic ecosystem of competing agents as a viable path to steer them in more human-aligned trends and mitigate risks. We explore how misalignment may serve and should be promoted as a counterbalancing mechanism to team up with whichever agents are most aligned to human interests, ensuring that no single system dominates destructively. The main premise of our contribution is that misalignment is inevitable because full AI-human alignment is a mathematical impossibility from Turing-complete systems, which we also offer as a proof in this contribution, a feature then inherited to AGI and ASI systems. We introduce a change-of-opinion attack test based on perturbation and intervention analysis to study how humans and agents may change or neutralise friendly and unfriendly AIs through cooperation and competition. We show that open models are more diverse and that most likely guardrails implemented in proprietary models are successful at controlling some of the agents' range of behaviour with positive and negative consequences while closed systems are more steerable and can also be used against proprietary AI systems. We also show that human and AI intervention has different effects hence suggesting multiple strategies.","abs_url":"https://arxiv.org/abs/2505.02581","pdf_url":"https://arxiv.org/pdf/2505.02581v4","match":"abstract"},{"id":"2504.17404","version":5,"title":"Super Co-alignment of Human and AI for Sustainable Symbiotic Society","authors":["Yi Zeng","Feifei Zhao","Yuwei Wang","Enmeng Lu","Yaodong Yang","Lei Wang","Chao Liu","Yitao Liang","Dongcheng Zhao","Bing Han","Haibo Tong","Yao Liang","Dongqi Liang","Kang Sun","Boyuan Chen","Jinyu Fan"],"published":"2025-04-24","updated":"2025-06-28","primary_category":"cs.AI","categories":["cs.AI"],"abstract":"As Artificial Intelligence (AI) advances toward Artificial General Intelligence (AGI) and eventually Artificial Superintelligence (ASI), it may potentially surpass human control, deviate from human values, and even lead to irreversible catastrophic consequences in extreme cases. This looming risk underscores the critical importance of the \"superalignment\" problem - ensuring that AI systems which are much smarter than humans, remain aligned with human (compatible) intentions and values. While current scalable oversight and weak-to-strong generalization methods demonstrate certain applicability, they exhibit fundamental flaws in addressing the superalignment paradigm - notably, the unidirectional imposition of human values cannot accommodate superintelligence's autonomy or ensure AGI/ASI's stable learning. We contend that the values for sustainable symbiotic society should be co-shaped by humans and living AI together, achieving \"Super Co-alignment.\" Guided by this vision, we propose a concrete framework that integrates external oversight and intrinsic proactive alignment. External oversight superalignment should be grounded in human-centered ultimate decision, supplemented by interpretable automated evaluation and correction, to achieve continuous alignment with humanity's evolving values. Intrinsic proactive superalignment is rooted in a profound understanding of the Self, others, and society, integrating self-awareness, self-reflection, and empathy to spontaneously infer human intentions, distinguishing good from evil and proactively prioritizing human well-being. The integration of externally-driven oversight with intrinsically-driven proactive alignment will co-shape symbiotic values and rules through iterative human-ASI co-alignment, paving the way for achieving safe and beneficial AGI and ASI for good, for human, and for a symbiotic ecology.","abs_url":"https://arxiv.org/abs/2504.17404","pdf_url":"https://arxiv.org/pdf/2504.17404v5","match":"abstract"},{"id":"2504.05259","version":1,"title":"How to evaluate control measures for LLM agents? A trajectory from today to superintelligence","authors":["Tomek Korbak","Mikita Balesni","Buck Shlegeris","Geoffrey Irving"],"published":"2025-04-07","updated":"2025-04-07","primary_category":"cs.AI","categories":["cs.AI","cs.CR"],"abstract":"As LLM agents grow more capable of causing harm autonomously, AI developers will rely on increasingly sophisticated control measures to prevent possibly misaligned agents from causing harm. AI developers could demonstrate that their control measures are sufficient by running control evaluations: testing exercises in which a red team produces agents that try to subvert control measures. To ensure control evaluations accurately capture misalignment risks, the affordances granted to this red team should be adapted to the capability profiles of the agents to be deployed under control measures. In this paper we propose a systematic framework for adapting affordances of red teams to advancing AI capabilities. Rather than assuming that agents will always execute the best attack strategies known to humans, we demonstrate how knowledge of an agents's actual capability profile can inform proportional control evaluations, resulting in more practical and cost-effective control measures. We illustrate our framework by considering a sequence of five fictional models (M1-M5) with progressively advanced capabilities, defining five distinct AI control levels (ACLs). For each ACL, we provide example rules for control evaluation, control measures, and safety cases that could be appropriate. Finally, we show why constructing a compelling AI control safety case for superintelligent LLM agents will require research breakthroughs, highlighting that we might eventually need alternative approaches to mitigating misalignment risk.","abs_url":"https://arxiv.org/abs/2504.05259","pdf_url":"https://arxiv.org/pdf/2504.05259v1","match":"title"},{"title":"Superintelligence, heuristics and embodied threats","authors":["Antonio Mastrogiorgio","Riccardo Palumbo"],"published":"2025-03-28","abs_url":"https://link.springer.com/article/10.1007/s11299-025-00317-0","primary_category":"Mind & Society","abstract":"The superintelligence debate raises the issue of the existential risk for human civilization posed by superintelligent artificial systems, which, in the long term, could surpass human intelligence. This debate, focusing on the computational dimension of intelligence, overlooks the current trends in behavioral sciences and bio-inspired robotics, emphasizing that cognition is embodied, embedded, enactive, and extended, and works through interactions with the environment. In this contribution, we discuss the potential threat arising from a near-future generation of cheap-but-effective artificial systems that are computationally poor but can enact intelligent behaviors through their embodied interactions with the environment.","venue":true},{"title":"Olympians: humanity as a solution to the control problem for artificial superintelligence","authors":["Daniel McKay"],"published":"2025-03-19","abs_url":"https://link.springer.com/article/10.1007/s43681-025-00712-w","primary_category":"AI and Ethics","abstract":"The control problem for artificial superintelligences is both difficult to solve and highly costly to get wrong. In this paper, I outline the problems with current methods of solving this problem and pose a novel solution. I argue that by using a human mind as the basis for an artificial superintelligence, we can mitigate some of the dangers that such a superintelligence would pose. I call this type of human-based artificial superintelligence an Olympian.","venue":true},{"title":"On Artificial Superintelligence and the Problem of Charismatic Extinction Threats","authors":["Nicholas Agar","Murilo Vilaça"],"published":"2025-03-10","abs_url":"https://jeet.ieet.org/index.php/home/article/view/166","primary_category":"Journal of Ethics and Emerging Technologies","abstract":"This paper focuses on the challenge of finding a rational response to Artificial Superintelligence (ASI) as an extinction threat. We allow that artificially superintelligent beings are possible. This leaves open the question of how much we should worry about them. We treat ASI as a charismatic extinction threat. Our starting analysis of charisma comes from the sociologist Max Weber (1947). We extend the concept of charisma beyond individual personalities to events including extinction threats. Our principal example of a charismatic extinction threat is Skynet, the human-unfriendly AI of the movies of the Terminator franchise. Skynet’s charisma interferes with the processes by which we rationally evaluate future risks. Our exploration of the psychological and emotional dimensions of assessing extinction threats considers work by the Nobel laureate economist Robert Shiller (2019) in the emerging field of narrative economics. We connect the virality of extinction stories with the work of the psychologist Elke Weber. According to Elke Weber (2010) we have a finite pool of worry to allocate to all of our future concerns. The charisma of Skynet means that we risk worrying too much about it and, as a consequence, worrying insufficiently about the uncharismatic challenge of climate change. We conclude with a brief discussion of a proposal that could lead to a more rational allocation of worry about extinction and other threats to humanity. We counsel imagination insurance for an intrinsically uncertain future.","venue":true},{"id":"2503.07660","version":2,"title":"Research Superalignment Should Advance Now with Alternating Competence and Conformity Optimization","authors":["HyunJin Kim","Xiaoyuan Yi","Jing Yao","Muhua Huang","JinYeong Bak","James Evans","Xing Xie"],"published":"2025-03-07","updated":"2026-02-09","primary_category":"cs.AI","categories":["cs.AI","cs.CY","cs.LG"],"abstract":"The recent leap in AI capabilities, driven by big generative models, has sparked the possibility of achieving Artificial General Intelligence (AGI) and further triggered discussions on Artificial Superintelligence (ASI)-a system surpassing all humans across measured domains. This gives rise to the critical research question of: As we approach ASI, how do we align it with human values, ensuring it benefits rather than harms human society, a.k.a., the Superalignment problem. Despite ASI being regarded by many as a hypothetical concept, in this position paper, we argue that superalignment is achievable and research on it should advance immediately, through simultaneous and alternating optimization of task competence and value conformity. We posit that superalignment is not merely a safeguard for ASI but also necessary for its responsible realization. To support this position, we first provide a formal definition of superalignment rooted in the gap between capability and capacity, delve into its perceived infeasibility by analyzing the limitations of existing paradigms, and then illustrate a conceptual path of superalignment to support its achievability, centered on two fundamental principles. This work frames a potential initiative for developing value-aligned next-generation AI in the future, which will garner greater benefits and reduce potential harm to humanity.","abs_url":"https://arxiv.org/abs/2503.07660","pdf_url":"https://arxiv.org/pdf/2503.07660v2","match":"abstract"},{"id":"2503.05628","version":2,"title":"Superintelligence Strategy: Expert Version","authors":["Dan Hendrycks","Eric Schmidt","Alexandr Wang"],"published":"2025-03-07","updated":"2025-04-14","primary_category":"cs.CY","categories":["cs.CY","cs.AI"],"abstract":"Rapid advances in AI are beginning to reshape national security. Destabilizing AI developments could rupture the balance of power and raise the odds of great-power conflict, while widespread proliferation of capable AI hackers and virologists would lower barriers for rogue actors to cause catastrophe. Superintelligence -- AI vastly better than humans at nearly all cognitive tasks -- is now anticipated by AI researchers. Just as nations once developed nuclear strategies to secure their survival, we now need a coherent superintelligence strategy to navigate a new period of transformative change. We introduce the concept of Mutual Assured AI Malfunction (MAIM): a deterrence regime resembling nuclear mutual assured destruction (MAD) where any state's aggressive bid for unilateral AI dominance is met with preventive sabotage by rivals. Given the relative ease of sabotaging a destabilizing AI project -- through interventions ranging from covert cyberattacks to potential kinetic strikes on datacenters -- MAIM already describes the strategic picture AI superpowers find themselves in. Alongside this, states can increase their competitiveness by bolstering their economies and militaries through AI, and they can engage in nonproliferation to rogue actors to keep weaponizable AI capabilities out of their hands. Taken together, the three-part framework of deterrence, nonproliferation, and competitiveness outlines a robust strategy to superintelligence in the years ahead.","abs_url":"https://arxiv.org/abs/2503.05628","pdf_url":"https://arxiv.org/pdf/2503.05628v2","match":"both"},{"id":"2503.15508","version":1,"title":"Assessing Human Intelligence Augmentation Strategies Using Brain Machine Interfaces and Brain Organoids in the Era of AI Advancement","authors":["Kenta Kitamura"],"published":"2025-01-27","updated":"2025-01-27","primary_category":"cs.HC","categories":["cs.HC","cs.CR","cs.ET"],"abstract":"The rapid advancement of Artificial Intelligence (AI) technologies, including the potential emergence of Artificial General Intelligence (AGI) and Artificial Superintelligence (ASI), has raised concerns about AI surpassing human cognitive capabilities. To address this challenge, intelligence augmentation approaches, such as Brain Machine Interfaces (BMI) and Brain Organoid (BO) integration have been proposed. In this study, we compare three intelligence augmentation strategies, namely BMI, BO, and a hybrid approach combining both. These strategies are evaluated from three key perspectives that influence user decisions in selecting an augmentation method: information processing capacity, identity risk, and consent authenticity risk. First, we model these strategies and assess them across the three perspectives. The results reveal that while BO poses identity risks and BMI has limitations in consent authenticity capacity, the hybrid approach mitigates these weaknesses by striking a balance between the two. Second, we investigate how users might choose among these intelligence augmentation strategies in the context of evolving AI capabilities over time. As the result, we find that BMI augmentation alone is insufficient to compete with advanced AI, and while BO augmentation offers scalability, BO increases identity risks as the scale grows. Moreover, the hybrid approach provides a balanced solution by adapting to AI advancements. This study provides a novel framework for human capability augmentation in the era of advancing AI and serves as a guideline for adapting to AI development.","abs_url":"https://arxiv.org/abs/2503.15508","pdf_url":"https://arxiv.org/pdf/2503.15508v1","match":"abstract"},{"id":"2501.06948","version":1,"title":"The Einstein Test: Towards a Practical Test of a Machine's Ability to Exhibit Superintelligence","authors":["David Benrimoh","Nace Mikus","Ariel Rosenfeld"],"published":"2025-01-12","updated":"2025-01-12","primary_category":"cs.AI","categories":["cs.AI"],"abstract":"Creative and disruptive insights (CDIs), such as the development of the theory of relativity, have punctuated human history, marking pivotal shifts in our intellectual trajectory. Recent advancements in artificial intelligence (AI) have sparked debates over whether state of the art models possess the capacity to generate CDIs. We argue that the ability to create CDIs should be regarded as a significant feature of machine superintelligence (SI).To this end, we propose a practical test to evaluate whether an approach to AI targeting SI can yield novel insights of this kind. We propose the Einstein test: given the data available prior to the emergence of a known CDI, can an AI independently reproduce that insight (or one that is formally equivalent)? By achieving such a milestone, a machine can be considered to at least match humanity's past top intellectual achievements, and therefore to have the potential to surpass them.","abs_url":"https://arxiv.org/abs/2501.06948","pdf_url":"https://arxiv.org/pdf/2501.06948v1","match":"both"},{"title":"The need for an empirical research program regarding human–AI relational norms","authors":["Madeline G. Reinecke","Andreas Kappes","Sebastian Porsdam Mann","Julian Savulescu","Brian D. Earp"],"published":"2025-01-09","abs_url":"https://link.springer.com/article/10.1007/s43681-024-00631-2","primary_category":"AI and Ethics","abstract":"As artificial intelligence (AI) systems begin to take on social roles traditionally filled by humans, it will be crucial to understand how this affects people’s cooperative expectations. In the case of human–human dyads, different relationships are governed by different norms: For example, how two strangers—versus two friends or colleagues—should interact when faced with a similar coordination problem often differs. How will the rise of ‘social’ artificial intelligence (and ultimately, superintelligent AI) complicate people’s expectations about the cooperative norms that should govern different types of relationships, whether human–human or human–AI? Do people expect AI to adhere to the same cooperative dynamics as humans when in a given social role? Conversely, will they begin to expect humans in certain types of relationships to act more like AI? Here, we consider how people’s cooperative expectations may pull apart between human–human and human–AI relationships, detailing an empirical proposal for mapping these distinctions across relationship types. We see the data resulting from our proposal as relevant for understanding people’s relationship–specific cooperative expectations in an age of social AI, which may also forecast potential resistance towards AI systems occupying certain social roles. Finally, these data can form the basis for ethical evaluations: What relationship–specific cooperative norms we should adopt for human–AI interactions, or reinforce through responsible AI design, depends partly on empirical facts about what norms people find intuitive for such interactions (along with the costs and benefits of maintaining these). Toward the end of the paper, we discuss how these relational norms may change over time and consider the implications of this for the proposed research program.","venue":true},{"id":"2501.14749","version":1,"title":"The Manhattan Trap: Why a Race to Artificial Superintelligence is Self-Defeating","authors":["Corin Katzke","Gideon Futerman"],"published":"2024-12-22","updated":"2024-12-22","primary_category":"cs.CY","categories":["cs.CY"],"abstract":"This paper examines the strategic dynamics of international competition to develop Artificial Superintelligence (ASI). We argue that the same assumptions that might motivate the US to race to develop ASI also imply that such a race is extremely dangerous. These assumptions--that ASI would provide a decisive military advantage and that states are rational actors prioritizing survival--imply that a race would heighten three critical risks: great power conflict, loss of control of ASI systems, and the undermining of liberal democracy. Our analysis shows that ASI presents a trust dilemma rather than a prisoners dilemma, suggesting that international cooperation to control ASI development is both preferable and strategically sound. We conclude that cooperation is achievable.","abs_url":"https://arxiv.org/abs/2501.14749","pdf_url":"https://arxiv.org/pdf/2501.14749v1","match":"both"},{"id":"2412.16468","version":4,"title":"The Road to Artificial SuperIntelligence: A Comprehensive Survey of Superalignment","authors":["HyunJin Kim","DongHyun Ryu","Xiaoyuan Yi","Jing Yao","Jianxun Lian","Muhua Huang","Shitong Duan","JinYeong Bak","Xing Xie"],"published":"2024-12-20","updated":"2026-06-17","primary_category":"cs.LG","categories":["cs.LG"],"abstract":"The emergence of large language models (LLMs) has sparked discussion on Artificial Superintelligence (ASI), a hypothetical AI system that surpasses human intelligence. Although ASI remains hypothetical and far beyond current AI capabilities, discussing its potential and exploring its feasibility and potential risks is critical for the development of future AI systems. The idea of superalignment originates from scalable oversight, which studies how to supervise increasingly capable AI systems when direct human supervision becomes insufficient. In this paper, we focus on the superalignment problem: \"The process of supervising, controlling, and governing artificial superintelligence.\" We first review scalable oversight paradigms-Sandwiching, Self-Enhancement, and Weak-to-Strong Generalization -- then analyze the limitations of current paradigms through the lens of possibility and impossibility, discuss key challenges, and propose pathways for the safe and continual improvement of future AI systems.","abs_url":"https://arxiv.org/abs/2412.16468","pdf_url":"https://arxiv.org/pdf/2412.16468v4","match":"both"},{"title":"Why AI may undermine phronesis and what to do about it","authors":["Cheng-hung Tsai","Hsiu-lin Ku"],"published":"2024-12-16","abs_url":"https://link.springer.com/article/10.1007/s43681-024-00617-0","primary_category":"AI and Ethics","abstract":"Phronesis, or practical wisdom, is a capacity the possession of which enables one to make good practical judgments and thus fulfill the distinctive function of human beings. Nir Eisikovits and Dan Feldman convincingly argue that this capacity may be undermined by statistical machine-learning-based AI. The critic questions: why should we worry that AI undermines phronesis? Why can’t we epistemically defer to AI, especially when it is superintelligent? Eisikovits and Feldman acknowledge such objection but do not consider it seriously. In this paper, we argue that there is a way to reconcile Eisikovits and Feldman with their critic by adopting the principle of epistemic heed, according to which we should exercise our rational capacity as much as possible while heeding a superintelligence’s output whenever possible.","venue":true},{"id":"2412.07278","version":1,"title":"Superficial Consciousness Hypothesis for Autoregressive Transformers","authors":["Yosuke Miyanishi","Keita Mitani"],"published":"2024-12-10","updated":"2024-12-10","primary_category":"cs.AI","categories":["cs.AI","cs.IT"],"abstract":"The alignment between human objectives and machine learning models built on these objectives is a crucial yet challenging problem for achieving Trustworthy AI, particularly when preparing for superintelligence (SI). First, given that SI does not exist today, empirical analysis for direct evidence is difficult. Second, SI is assumed to be more intelligent than humans, capable of deceiving us into underestimating its intelligence, making output-based analysis unreliable. Lastly, what kind of unexpected property SI might have is still unclear. To address these challenges, we propose the Superficial Consciousness Hypothesis under Information Integration Theory (IIT), suggesting that SI could exhibit a complex information-theoretic state like a conscious agent while unconscious. To validate this, we use a hypothetical scenario where SI can update its parameters \"at will\" to achieve its own objective (mesa-objective) under the constraint of the human objective (base objective). We show that a practical estimate of IIT's consciousness metric is relevant to the widely used perplexity metric, and train GPT-2 with those two objectives. Our preliminary result suggests that this SI-simulating GPT-2 could simultaneously follow the two objectives, supporting the feasibility of the Superficial Consciousness Hypothesis.","abs_url":"https://arxiv.org/abs/2412.07278","pdf_url":"https://arxiv.org/pdf/2412.07278v1","match":"abstract"},{"title":"Signs of consciousness in AI: Can GPT-3 tell how smart it really is?","authors":["Ljubiša Bojić","Irena Stojković","Zorana Jolić Marjanović"],"published":"2024-12-02","abs_url":"https://www.nature.com/articles/s41599-024-04154-3","primary_category":"Humanities and Social Sciences Communications","abstract":"The emergence of artificial intelligence (AI) is transforming how humans live and interact, raising both excitement and concerns—particularly about the potential for AI consciousness. For example, Google engineer Blake Lemoine suggested that the AI chatbot LaMDA might become sentient. At that time, GPT-3 was one of the most powerful publicly available language models, capable of simulating human reasoning to a certain extent. The notion of GPT-3 having some degree of consciousness could be linked to its ability to produce human-like responses, hinting at a basic level of understanding. To explore this further, we administered both objective and self-assessment tests of cognitive (CI) and emotional intelligence (EI) to GPT-3. Results showed that GPT-3 outperformed average humans on CI tests requiring the use and demonstration of acquired knowledge. However, its logical reasoning and EI capacities matched those of an average human. GPT-3’s self-assessments of CI and EI didn’t always align with its objective performance, with variations comparable to different human subsamples (e.g., high performers, males). A further discussion considered whether these results signal emerging subjectivity and self-awareness in AI. Future research should examine various language models to identify emergent properties of AI. The goal is not to discover machine consciousness itself, but to identify signs of its development, occurring independently of training and fine-tuning processes. If AI is to be further developed and widely deployed in human interactions, creating empathic AI that mimics human behavior is essential. The rapid advancement toward superintelligence requires continuous monitoring of AI’s human-like capabilities, particularly in general-purpose models, to ensure safety and alignment with human values.","venue":true},{"title":"Aesthetic Value and the AI Alignment Problem","authors":["Alice C. Helliwell"],"published":"2024-11-05","abs_url":"https://link.springer.com/article/10.1007/s13347-024-00816-x","primary_category":"Philosophy & Technology","abstract":"The threat from possible future superintelligent AI has given rise to discussion of the so-called “value alignment problem”. This is the problem of how to ensure artificially intelligent systems align with human values, and thus (hopefully) mitigate risks associated with them. Naturally, AI value alignment is often discussed in relation to morally relevant values, such as the value of human lives or human wellbeing. However, solutions to the value alignment problem target all human values, not only morally relevant ones. Is there a value alignment problem in other domains? In this paper, I explore whether the AI value alignment problem extends beyond morally relevant values to include aesthetic values. I demonstrate that the value alignment problem as typically framed includes aesthetic values, and, using examples from computer vision, put forward that AI may be misaligned with human values in the artistic realm. Whilst misalignment may be cause for concern when considering AI creativity, I argue that aesthetic value is a case in which we do not want AI to be fully aligned with human values. In doing so, I offer support to Peterson’s moderate value alignment thesis.","venue":true},{"id":"2410.19308","version":1,"title":"Semantics in Robotics: Environmental Data Can't Yield Conventions of Human Behaviour","authors":["Jamie Milton Freestone"],"published":"2024-10-25","updated":"2024-10-25","primary_category":"cs.RO","categories":["cs.RO","cs.AI"],"abstract":"The word semantics, in robotics and AI, has no canonical definition. It usually serves to denote additional data provided to autonomous agents to aid HRI. Most researchers seem, implicitly, to understand that such data cannot simply be extracted from environmental data. I try to make explicit why this is so and argue that so-called semantics are best understood as data comprised of conventions of human behaviour. This includes labels, most obviously, but also places, ontologies, and affordances. Object affordances are especially problematic because they require not only semantics that are not in the environmental data (conventions of object use) but also an understanding of physics and object combinations that would, if achieved, constitute artificial superintelligence.","abs_url":"https://arxiv.org/abs/2410.19308","pdf_url":"https://arxiv.org/pdf/2410.19308v1","match":"abstract"},{"title":"The selfish machine? On the power and limitation of natural selection to understand the development of advanced AI","authors":["Maarten Boudry","Simon Friederich"],"published":"2024-09-24","abs_url":"https://link.springer.com/article/10.1007/s11098-024-02226-3","primary_category":"Philosophical Studies","abstract":"Some philosophers and machine learning experts have speculated that superintelligent Artificial Intelligences (AIs), if and when they arrive on the scene, will wrestle away power from humans, with potentially catastrophic consequences. Dan Hendrycks has recently buttressed such worries by arguing that AI systems will undergo evolution by natural selection, which will endow them with instinctive drives for self-preservation, dominance and resource accumulation that are typical of evolved creatures. In this paper, we argue that this argument is not compelling as it stands. Evolutionary processes, as we point out, can be more or less Darwinian along a number of dimensions. Making use of Peter Godfrey-Smith’s framework of Darwinian spaces, we argue that the more evolution is top-down, directed and driven by intelligent agency, the less paradigmatically Darwinian it becomes. We then apply the concept of “domestication” to AI evolution, which, although theoretically satisfying the minimal definition of natural selection, is channeled through the minds of fore-sighted and intelligent agents, based on selection criteria desirable to them (which could be traits like docility, obedience and non-aggression). In the presence of such intelligent planning, it is not clear that selection of AIs, even selection in a competitive and ruthless market environment, will end up favoring “selfish” traits. In the end, however, we do agree with Hendrycks’ conditionally: If superintelligent AIs end up “going feral” and competing in a truly Darwinian fashion, reproducing autonomously and without human supervision, this could pose a grave danger to human societies.","venue":true},{"title":"Is superintelligence necessarily moral?","authors":["Leonard Dung"],"published":"2024-09-24","abs_url":"https://academic.oup.com/analysis/article/84/4/730/7774058","primary_category":"Analysis 84(4): 730–738","abstract":"Numerous authors have expressed concern that advanced artificial intelligence (AI) poses an existential risk to humanity. These authors argue that we might build AI which is vastly intellectually superior to humans (a ‘superintelligence’), and which optimizes for goals that strike us as morally bad, or even irrational. Thus this argument assumes that a superintelligence might have morally bad goals. However, according to some views, a superintelligence necessarily has morally adequate goals. This might be the case either because abilities for moral reasoning and intelligence mutually depend on each other, or because moral realism and moral internalism are true. I argue that the former argument misconstrues the view that intelligence and goals are independent, and that the latter argument misunderstands the implications of moral internalism. Moreover, the current state of AI research provides additional reasons to think that a superintelligence could have bad goals.","venue":true},{"title":"Artificial consciousness in AI: a posthuman fallacy","authors":["M. Prabhu","J. Anil Premraj"],"published":"2024-09-14","abs_url":"https://link.springer.com/article/10.1007/s00146-024-02061-4","primary_category":"AI & SOCIETY","abstract":"Obsession toward technology has a long background of parallel evolution between humans and machines. This obsession became irrevocable when AI began to be a part of our daily lives. However, this AI integration became a subject of controversy when the fear of AI advancement in acquiring consciousness crept among mankind. Artificial consciousness is a long-debated topic in the field of artificial intelligence and neuroscience which has many ethical challenges and threats in society ranging from daily chores to Mars missions. This paper deals with the impact of AI-based science fiction films in society. This study aims to investigate the fascinating AI concept of artificial consciousness in light of posthuman terminology, technological singularity and superintelligence by analyzing the set of science fiction films to project the actual difference between science fictional AI and operational AI. Further, this paper explores the theoretical possibilities of artificial consciousness through a range of neuroscientific theories that are related to AI development. These theories are built toward prospective artificial consciousness in AI. This study discloses the posthuman fallacies that are built around the fear of AI acquiring artificial consciousness and its outcome.","venue":true},{"title":"The state as a model for AI control and alignment","authors":["Micha Elsner"],"published":"2024-09-11","abs_url":"https://link.springer.com/article/10.1007/s00146-024-02063-2","primary_category":"AI & SOCIETY","abstract":"Debates about the development of artificial superintelligence and its potential threats to humanity tend to assume that such a system would be historically unprecedented, and that its behavior must be predicted from first principles. I argue that this is not true: we can analyze multiagent intelligent systems (the best candidates for practical superintelligence) by comparing them to states, which also unite heterogeneous intelligences to achieve superhuman goals. States provide a model for several problems discussed in the literature on superintelligence, such as principal-agent problems and Instrumental Convergence. Philosophical arguments about governance, therefore, provide possible solutions to these problems, or point out problems in previously suggested solutions. In particular, the liberal concept of checks and balances, and Hannah Arendt’s concept of legitimacy, describe how state behavior is constrained by the preferences of constituents that could also apply to artificial systems. However, they also point out ways in which present-day computational developments could destabilize the international order by reducing the number of decision-makers involved in state actions. Thus, interstate competition not only serves as a model for the behavior of dangerous computational intelligences but also as the impetus for their development.","venue":true},{"id":"2407.20208","version":3,"title":"Supertrust foundational alignment: mutual trust must replace permanent control for safe superintelligence","authors":["James M. Mazzu"],"published":"2024-07-29","updated":"2024-11-28","primary_category":"cs.AI","categories":["cs.AI","cs.LG","cs.NE"],"abstract":"It's widely expected that humanity will someday create AI systems vastly more intelligent than us, leading to the unsolved alignment problem of \"how to control superintelligence.\" However, this commonly expressed problem is not only self-contradictory and likely unsolvable, but current strategies to ensure permanent control effectively guarantee that superintelligent AI will distrust humanity and consider us a threat. Such dangerous representations, already embedded in current models, will inevitably lead to an adversarial relationship and may even trigger the extinction event many fear. As AI leaders continue to \"raise the alarm\" about uncontrollable AI, further embedding concerns about it \"getting out of our control\" or \"going rogue,\" we're unintentionally reinforcing our threat and deepening the risks we face. The rational path forward is to strategically replace intended permanent control with intrinsic mutual trust at the foundational level. The proposed Supertrust alignment meta-strategy seeks to accomplish this by modeling instinctive familial trust, representing superintelligence as the evolutionary child of human intelligence, and implementing temporary controls/constraints in the manner of effective parenting. Essentially, we're creating a superintelligent \"child\" that will be exponentially smarter and eventually independent of our control. We therefore have a critical choice: continue our controlling intentions and usher in a brief period of dominance followed by extreme hardship for humanity, or intentionally create the foundational mutual trust required for long-term safe coexistence.","abs_url":"https://arxiv.org/abs/2407.20208","pdf_url":"https://arxiv.org/pdf/2407.20208v3","match":"both"},{"title":"A Collective Intelligence Approach to Safe Artificial General Intelligence","authors":["Craig A. Kaplan"],"published":"2024-07-17","abs_url":"https://link.springer.com/chapter/10.1007/978-3-031-65572-2_12","primary_category":"International Conference on Artificial General Intelligence — Artificial General Intelligence (LNCS)","abstract":"If Artificial General Intelligence (AGI) proves to be a “winner-take-all” scenario where the first company or country to develop AGI dominates, then the first AGI must also be the safest. The safest, and fastest, path to AGI may be to harness the collective intelligence of multiple AI and human agents in an AGI network. This approach has roots in seminal ideas from four of the scientists who founded the field of AI: Allen Newell, Marvin Minsky, Claude Shannon, and Herbert Simon. Extrapolating key insights and combining them with the work of modern researchers, illuminates a fast and safe path to AGI. The seminal ideas discussed are 1) Society of Mind (Minsky), 2) Information Theory (Shannon), 3) Problem Solving Theory (Newell & Simon), and 4) Bounded Rationality (Simon). Society of Mind describes a collective intelligence approach that can be used with AI and human agents to create an AGI network. Information Theory helps address the critical issue of how an AGI system will increase its intelligence over time. Problem Solving Theory provides a universal framework that AI and human agents can use to communicate efficiently, effectively, and safely. Bounded Rationality helps us better understand not only the capabilities of SuperIntelligent AGI but also how humans can remain relevant where the intelligence of AGI vastly exceeds that of its human creators. Each key idea can be combined with recent work in the fields of Artificial Intelligence, Machine Learning, and Large Language Models to accelerate the development of a working, safe, AGI system.","venue":true},{"id":"2406.16772","version":2,"title":"OlympicArena Medal Ranks: Who Is the Most Intelligent AI So Far?","authors":["Zhen Huang","Zengzhi Wang","Shijie Xia","Pengfei Liu"],"published":"2024-06-24","updated":"2024-06-26","primary_category":"cs.CL","categories":["cs.CL","cs.AI"],"abstract":"In this report, we pose the following question: Who is the most intelligent AI model to date, as measured by the OlympicArena (an Olympic-level, multi-discipline, multi-modal benchmark for superintelligent AI)? We specifically focus on the most recently released models: Claude-3.5-Sonnet, Gemini-1.5-Pro, and GPT-4o. For the first time, we propose using an Olympic medal Table approach to rank AI models based on their comprehensive performance across various disciplines. Empirical results reveal: (1) Claude-3.5-Sonnet shows highly competitive overall performance over GPT-4o, even surpassing GPT-4o on a few subjects (i.e., Physics, Chemistry, and Biology). (2) Gemini-1.5-Pro and GPT-4V are ranked consecutively just behind GPT-4o and Claude-3.5-Sonnet, but with a clear performance gap between them. (3) The performance of AI models from the open-source community significantly lags behind these proprietary models. (4) The performance of these models on this benchmark has been less than satisfactory, indicating that we still have a long way to go before achieving superintelligence. We remain committed to continuously tracking and evaluating the performance of the latest powerful models on this benchmark (available at https://github.com/GAIR-NLP/OlympicArena).","abs_url":"https://arxiv.org/abs/2406.16772","pdf_url":"https://arxiv.org/pdf/2406.16772v2","match":"abstract"},{"id":"2406.12753","version":2,"title":"OlympicArena: Benchmarking Multi-discipline Cognitive Reasoning for Superintelligent AI","authors":["Zhen Huang","Zengzhi Wang","Shijie Xia","Xuefeng Li","Haoyang Zou","Ruijie Xu","Run-Ze Fan","Lyumanshan Ye","Ethan Chern","Yixin Ye","Yikai Zhang","Yuqing Yang","Ting Wu","Binjie Wang","Shichao Sun","Yang Xiao","Yiyuan Li","Fan Zhou","Steffi Chern","Yiwei Qin","Yan Ma","Jiadi Su","Yixiu Liu","Yuxiang Zheng","Shaoting Zhang"],"published":"2024-06-18","updated":"2025-03-06","primary_category":"cs.CL","categories":["cs.CL","cs.AI"],"abstract":"The evolution of Artificial Intelligence (AI) has been significantly accelerated by advancements in Large Language Models (LLMs) and Large Multimodal Models (LMMs), gradually showcasing potential cognitive reasoning abilities in problem-solving and scientific discovery (i.e., AI4Science) once exclusive to human intellect. To comprehensively evaluate current models' performance in cognitive reasoning abilities, we introduce OlympicArena, which includes 11,163 bilingual problems across both text-only and interleaved text-image modalities. These challenges encompass a wide range of disciplines spanning seven fields and 62 international Olympic competitions, rigorously examined for data leakage. We argue that the challenges in Olympic competition problems are ideal for evaluating AI's cognitive reasoning due to their complexity and interdisciplinary nature, which are essential for tackling complex scientific challenges and facilitating discoveries. Beyond evaluating performance across various disciplines using answer-only criteria, we conduct detailed experiments and analyses from multiple perspectives. We delve into the models' cognitive reasoning abilities, their performance across different modalities, and their outcomes in process-level evaluations, which are vital for tasks requiring complex reasoning with lengthy solutions. Our extensive evaluations reveal that even advanced models like GPT-4o only achieve a 39.97% overall accuracy, illustrating current AI limitations in complex reasoning and multimodal integration. Through the OlympicArena, we aim to advance AI towards superintelligence, equipping it to address more complex challenges in science and beyond. We also provide a comprehensive set of resources to support AI research, including a benchmark dataset, an open-source annotation platform, a detailed evaluation tool, and a leaderboard with automatic submission features.","abs_url":"https://arxiv.org/abs/2406.12753","pdf_url":"https://arxiv.org/pdf/2406.12753v2","match":"abstract"},{"title":"On quantum computing for artificial superintelligence","authors":["Anna Grabowska","Artur Gunia"],"published":"2024-06-04","abs_url":"https://link.springer.com/article/10.1007/s13194-024-00584-7","primary_category":"European Journal for Philosophy of Science","abstract":"Artificial intelligence algorithms, fueled by continuous technological development and increased computing power, have proven effective across a variety of tasks. Concurrently, quantum computers have shown promise in solving problems beyond the reach of classical computers. These advancements have contributed to a misconception that quantum computers enable hypercomputation, sparking speculation about quantum supremacy leading to an intelligence explosion and the creation of superintelligent agents. We challenge this notion, arguing that current evidence does not support the idea that quantum technologies enable hypercomputation. Fundamental limitations on information storage within finite spaces and the accessibility of information from quantum states constrain quantum computers from surpassing the Turing computing barrier. While quantum technologies may offer exponential speed-ups in specific computing cases, there is insufficient evidence to suggest that focusing solely on quantum-related problems will lead to technological singularity and the emergence of superintelligence. Subsequently, there is no premise suggesting that general intelligence depends on quantum effects or that accelerating existing algorithms through quantum means will replicate true intelligence. We propose that if superintelligence is to be achieved, it will not be solely through quantum technologies. Instead, the attainment of superintelligence remains a conceptual challenge that humanity has yet to overcome, with quantum technologies showing no clear path toward its resolution.","venue":true},{"id":"2404.14387","version":2,"title":"A Survey on Self-Evolution of Large Language Models","authors":["Zhengwei Tao","Ting-En Lin","Xiancai Chen","Hangyu Li","Yuchuan Wu","Yongbin Li","Zhi Jin","Fei Huang","Dacheng Tao","Jingren Zhou"],"published":"2024-04-22","updated":"2024-06-03","primary_category":"cs.CL","categories":["cs.CL","cs.AI"],"abstract":"Large language models (LLMs) have significantly advanced in various fields and intelligent agent applications. However, current LLMs that learn from human or external model supervision are costly and may face performance ceilings as task complexity and diversity increase. To address this issue, self-evolution approaches that enable LLM to autonomously acquire, refine, and learn from experiences generated by the model itself are rapidly growing. This new training paradigm inspired by the human experiential learning process offers the potential to scale LLMs towards superintelligence. In this work, we present a comprehensive survey of self-evolution approaches in LLMs. We first propose a conceptual framework for self-evolution and outline the evolving process as iterative cycles composed of four phases: experience acquisition, experience refinement, updating, and evaluation. Second, we categorize the evolution objectives of LLMs and LLM-based agents; then, we summarize the literature and provide taxonomy and insights for each module. Lastly, we pinpoint existing challenges and propose future directions to improve self-evolution frameworks, equipping researchers with critical insights to fast-track the development of self-evolving LLMs. Our corresponding GitHub repository is available at https://github.com/AlibabaResearch/DAMO-ConvAI/tree/main/Awesome-Self-Evolution-of-LLM","abs_url":"https://arxiv.org/abs/2404.14387","pdf_url":"https://arxiv.org/pdf/2404.14387v2","match":"abstract"},{"title":"Instrumental divergence","authors":["J. Dmitri Gallow"],"published":"2024-04-06","abs_url":"https://link.springer.com/article/10.1007/s11098-024-02129-3","primary_category":"Philosophical Studies","abstract":"The thesis of instrumental convergence holds that a wide range of ends have common means: for instance, self preservation, desire preservation, self improvement, and resource acquisition. Bostrom contends that instrumental convergence gives us reason to think that “the default outcome of the creation of machine superintelligensome of the ‘convergence is existential catastrophe”. I use the tools of decision theory to investigate whether this thesis is true. I find that, even if intrinsic desires are randomly selected, instrumental rationality induces biases towards certain kinds of choices. Firstly, a bias towards choices which leave less up to chance. Secondly, a bias towards desire preservation, in line with Bostrom’s conjecture. And thirdly, a bias towards choices which afford more choices later on. I do not find biases towards any other of the convergent instrumental means on Bostrom’s list. I conclude that the biases induced by instrumental rationality at best weakly support Bostrom’s conclusion that machine superintelligence is likely to lead to existential catastrophe.","venue":true},{"id":"2405.00042","version":1,"title":"Is Artificial Intelligence the great filter that makes advanced technical civilisations rare in the universe?","authors":["Michael Garrett"],"published":"2024-04-01","updated":"2024-04-01","primary_category":"physics.pop-ph","categories":["physics.pop-ph","physics.soc-ph"],"abstract":"This study examines the hypothesis that the rapid development of Artificial Intelligence (AI), culminating in the emergence of Artificial Superintelligence (ASI), could act as a \"Great Filter\" that is responsible for the scarcity of advanced technological civilisations in the universe. It is proposed that such a filter emerges before these civilisations can develop a stable, multiplanetary existence, suggesting the typical longevity (L) of a technical civilization is less than 200 years. Such estimates for L, when applied to optimistic versions of the Drake equation, are consistent with the null results obtained by recent SETI surveys, and other efforts to detect various technosignatures across the electromagnetic spectrum. Through the lens of SETI, we reflect on humanity's current technological trajectory - the modest projections for L suggested here, underscore the critical need to quickly establish regulatory frameworks for AI development on Earth and the advancement of a multiplanetary society to mitigate against such existential threats. The persistence of intelligent and conscious life in the universe could hinge on the timely and effective implementation of such international regulatory measures and technological endeavours.","abs_url":"https://arxiv.org/abs/2405.00042","pdf_url":"https://arxiv.org/pdf/2405.00042v1","match":"abstract"},{"id":"2406.08492","version":1,"title":"ASI as the New God: Technocratic Theocracy","authors":["Tevfik Uyar"],"published":"2024-03-22","updated":"2024-03-22","primary_category":"cs.AI","categories":["cs.AI","cs.CY"],"abstract":"As Artificial General Intelligence edges closer to reality, Artificial Superintelligence does too. This paper argues that ASI's unparalleled capabilities might lead people to attribute godlike infallibility to it, resulting in a cognitive bias toward unquestioning acceptance of its decisions. By drawing parallels between ASI and divine attributes such as omnipotence, omniscience, and omnipresence, this analysis highlights the risks of conflating technological advancement with moral and ethical superiority. Such dynamics could engender a technocratic theocracy, where decision-making is abdicated to ASI, undermining human agency and critical thinking.","abs_url":"https://arxiv.org/abs/2406.08492","pdf_url":"https://arxiv.org/pdf/2406.08492v1","match":"abstract"},{"id":"2403.14681","version":1,"title":"AI Ethics: A Bibliometric Analysis, Critical Issues, and Key Gaps","authors":["Di Kevin Gao","Andrew Haverly","Sudip Mittal","Jiming Wu","Jingdao Chen"],"published":"2024-03-12","updated":"2024-03-12","primary_category":"cs.CY","categories":["cs.CY","cs.AI"],"abstract":"Artificial intelligence (AI) ethics has emerged as a burgeoning yet pivotal area of scholarly research. This study conducts a comprehensive bibliometric analysis of the AI ethics literature over the past two decades. The analysis reveals a discernible tripartite progression, characterized by an incubation phase, followed by a subsequent phase focused on imbuing AI with human-like attributes, culminating in a third phase emphasizing the development of human-centric AI systems. After that, they present seven key AI ethics issues, encompassing the Collingridge dilemma, the AI status debate, challenges associated with AI transparency and explainability, privacy protection complications, considerations of justice and fairness, concerns about algocracy and human enfeeblement, and the issue of superintelligence. Finally, they identify two notable research gaps in AI ethics regarding the large ethics model (LEM) and AI identification and extend an invitation for further scholarly research.","abs_url":"https://arxiv.org/abs/2403.14681","pdf_url":"https://arxiv.org/pdf/2403.14681v1","match":"abstract"},{"id":"2402.00667","version":1,"title":"Improving Weak-to-Strong Generalization with Scalable Oversight and Ensemble Learning","authors":["Jitao Sang","Yuhang Wang","Jing Zhang","Yanxu Zhu","Chao Kong","Junhong Ye","Shuyu Wei","Jinlin Xiao"],"published":"2024-02-01","updated":"2024-02-01","primary_category":"cs.CL","categories":["cs.CL"],"abstract":"This paper presents a follow-up study to OpenAI's recent superalignment work on Weak-to-Strong Generalization (W2SG). Superalignment focuses on ensuring that high-level AI systems remain consistent with human values and intentions when dealing with complex, high-risk tasks. The W2SG framework has opened new possibilities for empirical research in this evolving field. Our study simulates two phases of superalignment under the W2SG framework: the development of general superhuman models and the progression towards superintelligence. In the first phase, based on human supervision, the quality of weak supervision is enhanced through a combination of scalable oversight and ensemble learning, reducing the capability gap between weak teachers and strong students. In the second phase, an automatic alignment evaluator is employed as the weak supervisor. By recursively updating this auto aligner, the capabilities of the weak teacher models are synchronously enhanced, achieving weak-to-strong supervision over stronger student models.We also provide an initial validation of the proposed approach for the first phase. Using the SciQ task as example, we explore ensemble learning for weak teacher models through bagging and boosting. Scalable oversight is explored through two auxiliary settings: human-AI interaction and AI-AI debate. Additionally, the paper discusses the impact of improved weak supervision on enhancing weak-to-strong generalization based on in-context learning. Experiment code and dataset will be released at https://github.com/ADaM-BJTU/W2SG.","abs_url":"https://arxiv.org/abs/2402.00667","pdf_url":"https://arxiv.org/pdf/2402.00667v1","match":"abstract"},{"id":"2401.15109","version":1,"title":"Towards Collective Superintelligence: Amplifying Group IQ using Conversational Swarms","authors":["Louis Rosenberg","Gregg Willcox","Hans Schumann","Ganesh Mani"],"published":"2024-01-25","updated":"2024-01-25","primary_category":"cs.HC","categories":["cs.HC","cs.AI"],"abstract":"Swarm Intelligence (SI) is a natural phenomenon that enables biological groups to amplify their combined intellect by forming real-time systems. Artificial Swarm Intelligence (or Swarm AI) is a technology that enables networked human groups to amplify their combined intelligence by forming similar systems. In the past, swarm-based methods were constrained to narrowly defined tasks like probabilistic forecasting and multiple-choice decision making. A new technology called Conversational Swarm Intelligence (CSI) was developed in 2023 that amplifies the decision-making accuracy of networked human groups through natural conversational deliberations. The current study evaluated the ability of real-time groups using a CSI platform to take a common IQ test known as Raven's Advanced Progressive Matrices (RAPM). First, a baseline group of participants took the Raven's IQ test by traditional survey. This group averaged 45.6% correct. Then, groups of approximately 35 individuals answered IQ test questions together using a CSI platform called Thinkscape. These groups averaged 80.5% correct. This places the CSI groups in the 97th percentile of IQ test-takers and corresponds to an effective IQ increase of 28 points (p<0.001). This is an encouraging result and suggests that CSI is a powerful method for enabling conversational collective intelligence in large, networked groups. In addition, because CSI is scalable across groups of potentially any size, this technology may provide a viable pathway to building a Collective Superintelligence.","abs_url":"https://arxiv.org/abs/2401.15109","pdf_url":"https://arxiv.org/pdf/2401.15109v1","match":"both"},{"id":"2401.07836","version":3,"title":"Two Types of AI Existential Risk: Decisive and Accumulative","authors":["Atoosa Kasirzadeh"],"published":"2024-01-15","updated":"2025-01-17","primary_category":"cs.CY","categories":["cs.CY","cs.AI","cs.LG"],"abstract":"The conventional discourse on existential risks (x-risks) from AI typically focuses on abrupt, dire events caused by advanced AI systems, particularly those that might achieve or surpass human-level intelligence. These events have severe consequences that either lead to human extinction or irreversibly cripple human civilization to a point beyond recovery. This discourse, however, often neglects the serious possibility of AI x-risks manifesting incrementally through a series of smaller yet interconnected disruptions, gradually crossing critical thresholds over time. This paper contrasts the conventional \"decisive AI x-risk hypothesis\" with an \"accumulative AI x-risk hypothesis.\" While the former envisions an overt AI takeover pathway, characterized by scenarios like uncontrollable superintelligence, the latter suggests a different causal pathway to existential catastrophes. This involves a gradual accumulation of critical AI-induced threats such as severe vulnerabilities and systemic erosion of economic and political structures. The accumulative hypothesis suggests a boiling frog scenario where incremental AI risks slowly converge, undermining societal resilience until a triggering event results in irreversible collapse. Through systems analysis, this paper examines the distinct assumptions differentiating these two hypotheses. It is then argued that the accumulative view can reconcile seemingly incompatible perspectives on AI risks. The implications of differentiating between these causal pathways -- the decisive and the accumulative -- for the governance of AI as well as long-term AI safety are discussed.","abs_url":"https://arxiv.org/abs/2401.07836","pdf_url":"https://arxiv.org/pdf/2401.07836v3","match":"abstract"},{"id":"2402.00030","version":1,"title":"Evolution-Bootstrapped Simulation: Artificial or Human Intelligence: Which Came First?","authors":["Paul Alexander Bilokon"],"published":"2024-01-06","updated":"2024-01-06","primary_category":"cs.NE","categories":["cs.NE","cs.AI","q-bio.PE"],"abstract":"Humans have created artificial intelligence (AI), not the other way around. This statement is deceptively obvious. In this note, we decided to challenge this statement as a small, lighthearted Gedankenexperiment. We ask a simple question: in a world driven by evolution by natural selection, would neural networks or humans be likely to evolve first? We compare the Solomonoff--Kolmogorov--Chaitin complexity of the two and find neural networks (even LLMs) to be significantly simpler than humans. Further, we claim that it is unnecessary for any complex human-made equipment to exist for there to be neural networks. Neural networks may have evolved as naturally occurring objects before humans did as a form of chemical reaction-based or enzyme-based computation. Now that we know that neural networks can pass the Turing test and suspect that they may be capable of superintelligence, we ask whether the natural evolution of neural networks could lead from pure evolution by natural selection to what we call evolution-bootstrapped simulation. The evolution of neural networks does not involve irreducible complexity; would easily allow irreducible complexity to exist in the evolution-bootstrapped simulation; is a falsifiable scientific hypothesis; and is independent of / orthogonal to the issue of intelligent design.","abs_url":"https://arxiv.org/abs/2402.00030","pdf_url":"https://arxiv.org/pdf/2402.00030v1","match":"abstract"},{"title":"Tool, Teammate, Superintelligence: Identification of ChatGPT-Enabled Collaboration Patterns and their Benefits and Risks in Mutual Learning","authors":["Xusen Cheng","Shuang Zhang"],"published":"2024","abs_url":"https://scholarspace.manoa.hawaii.edu/items/edc40c14-b60a-495e-8059-177b9d04c435","primary_category":"Proceedings of the 57th Hawaii International Conference on System Sciences (HICSS 2024)","abstract":null,"venue":true},{"id":"2401.04112","version":1,"title":"Conversational Swarm Intelligence amplifies the accuracy of networked groupwise deliberations","authors":["Louis Rosenberg","Gregg Willcox","Hans Schumann","Ganesh Mani"],"published":"2023-12-19","updated":"2023-12-19","primary_category":"cs.HC","categories":["cs.HC"],"abstract":"Conversational Swarm Intelligence (CSI) is a communication technology that enables large, networked groups (25 to 2500 people) to hold real-time conversational deliberations online. Modeled on the dynamics of biological swarms, CSI enables the reasoning benefits of small-groups with the collective intelligence benefits of large-groups. In this pilot study, groups of 25 to 30 participants were asked to select players for a weekly Fantasy Football contest over an 11-week period. As a baseline, participants filled out a survey to record their player selections. As an experimental method, participants engaged in a real-time text-chat deliberation using a CSI platform called Thinkscape to collaboratively select sets of players. The results show that the real-time conversational group using CSI outperformed 66% of survey participants, demonstrating significant amplification of intelligence versus the median individual (p=0.020). The CSI method also significantly outperformed the most popular choices from the survey (the Wisdom of Crowd, p<0.001). These results suggest that CSI is an effective technology for amplifying the intelligence of groups engaged in real-time large-scale conversational deliberation and may offer a path to collective superintelligence.","abs_url":"https://arxiv.org/abs/2401.04112","pdf_url":"https://arxiv.org/pdf/2401.04112v1","match":"abstract"},{"title":"How Superintelligence Affects Human Health: A Scenario Analysis","authors":["Philipp Koebe","Tobias Schillings","Jan Oliver Schwarz"],"published":"2023-12","abs_url":"https://jfsdigital.org/articles-and-essays/2023-2/vol-28-no-2-december-2023/how-superintelligence-affects-human-health-a-scenario-analysis/","primary_category":"Journal of Futures Studies 28(2)","abstract":"Population health is a crucial determinant of human prosperity and well-being. Poor health can lead to reduced productivity, poverty, and premature death, with the COVID-19 pandemic underscoring the vulnerability of population health on a global scale. Self-learning algorithms have the potential to improve population health in a sustainable way and bring a paradigm shift to healthcare. We utilize intuitive logic to generate future scenarios in order to address the research question. These scenarios are categorized as either health-promoting or health-damaging, and superintelligence is considered either dominating or non-dominating. We provide strategic implications for each scenario, which can guide policy action in dealing with superintelligence.","venue":true},{"id":"2311.09452","version":4,"title":"Keep the Future Human: Why and How We Should Close the Gates to AGI and Superintelligence, and What We Should Build Instead","authors":["Anthony Aguirre"],"published":"2023-11-15","updated":"2025-03-07","primary_category":"cs.CY","categories":["cs.CY"],"abstract":"Dramatic advances in artificial intelligence over the past decade (for narrow-purpose AI) and the last several years (for general-purpose AI) have transformed AI from a niche academic field to the core business strategy of many of the world's largest companies, with hundreds of billions of dollars in annual investment in the techniques and technologies for advancing AI's capabilities. We now come to a critical juncture. As the capabilities of new AI systems begin to match and exceed those of humans across many cognitive domains, humanity must decide: how far do we go, and in what direction? This essay argues that we should keep the future human by closing the \"gates\" to smarter-than-human, autonomous, general-purpose AI -- sometimes called \"AGI\" -- and especially to the highly-superhuman version sometimes called \" superintelligence.\" Instead, we should focus on powerful, trustworthy AI tools that can empower individuals and transformatively improve human societies' abilities to do what they do best.","abs_url":"https://arxiv.org/abs/2311.09452","pdf_url":"https://arxiv.org/pdf/2311.09452v4","match":"both"},{"id":"2311.08706","version":1,"title":"Aligned: A Platform-based Process for Alignment","authors":["Ethan Shaotran","Ido Pesok","Sam Jones","Emi Liu"],"published":"2023-11-15","updated":"2023-11-15","primary_category":"cs.CY","categories":["cs.CY","cs.AI"],"abstract":"We are introducing Aligned, a platform for global governance and alignment of frontier models, and eventually superintelligence. While previous efforts at the major AI labs have attempted to gather inputs for alignment, these are often conducted behind closed doors. We aim to set the foundation for a more trustworthy, public-facing approach to safety: a constitutional committee framework. Initial tests with 680 participants result in a 30-guideline constitution with 93% overall support. We show the platform naturally scales, instilling confidence and enjoyment from the community. We invite other AI labs and teams to plug and play into the Aligned ecosystem.","abs_url":"https://arxiv.org/abs/2311.08706","pdf_url":"https://arxiv.org/pdf/2311.08706v1","match":"abstract"},{"id":"2311.00728","version":1,"title":"Towards Collective Superintelligence, a Pilot Study","authors":["Louis Rosenberg","Gregg Willcox","Hans Schumann"],"published":"2023-10-31","updated":"2023-10-31","primary_category":"cs.HC","categories":["cs.HC"],"abstract":"Conversational Swarm Intelligence (CSI) is a new technology that enables human groups of potentially any size to hold real-time deliberative conversations online. Modeled on the dynamics of biological swarms, CSI aims to optimize group insights and amplify group intelligence. It uses Large Language Models (LLMs) in a novel framework to structure large-scale conversations, combining the benefits of small-group deliberative reasoning and large-group collective intelligence. In this study, a group of 241 real-time participants were asked to estimate the number of gumballs in a jar by looking at a photo. In one test case, individual participants entered their estimation in a standard survey. In another test case, participants converged on groupwise estimates collaboratively using a prototype CSI text-chat platform called Thinkscape. The results show that when using CSI, the group of 241 participants estimated within 12% of the correct answer, which was significantly more accurate (p<0.001) than the average individual (mean error of 55%) and the survey-based Wisdom of Crowd (error of 25%). The group using CSI was also more accurate than an estimate generated by GPT 4 (error of 42%). This suggests that CSI is a viable method for enabling large, networked groups to hold coherent real-time deliberative conversations that amplify collective intelligence. Because this technology is scalable, it could provide a possible pathway towards building a general-purpose Collective Superintelligence (CSi).","abs_url":"https://arxiv.org/abs/2311.00728","pdf_url":"https://arxiv.org/pdf/2311.00728v1","match":"both"},{"title":"The energy challenges of artificial superintelligence","authors":["Klaus M. Stiefel","Jay S. Coggan"],"published":"2023-10-24","abs_url":"https://www.frontiersin.org/journals/artificial-intelligence/articles/10.3389/frai.2023.1240653/full","primary_category":"Frontiers in Artificial Intelligence","abstract":"We argue here that contemporary semiconductor computing technology poses a significant if not insurmountable barrier to the emergence of any artificial general intelligence system, let alone one anticipated by many to be \"superintelligent\". This limit on artificial superintelligence (ASI) emerges from the energy requirements of a system that would be more intelligent but orders of magnitude less efficient in energy use than human brains. An ASI would have to supersede not only a single brain but a large population given the effects of collective behavior on the advancement of societies, further multiplying the energy requirement. A hypothetical ASI would likely consume orders of magnitude more energy than what is available in highly industrialized nations. We estimate the energy use of ASI with an equation we term the \"Erasi equation\", for the Energy Requirement for Artificial SuperIntelligence. Additional efficiency consequences will emerge from the current unfocused and scattered developmental trajectory of AI research. Taken together, these arguments suggest that the emergence of an ASI is highly unlikely in the foreseeable future based on current computer architectures, primarily due to energy constraints, with biomimicry or other new technologies being possible solutions.","venue":true},{"title":"Answering Divine Love: Human Distinctiveness in the Light of Islam and Artificial Superintelligence","authors":["Yusuf Çelik"],"published":"2023-08-25","abs_url":"https://link.springer.com/article/10.1007/s11841-023-00977-w","primary_category":"Sophia","abstract":"In the Qur’an, human distinctiveness was first questioned by angels. These established denizens of the cosmos could not understand why God would create a seemingly pernicious human when immaculate devotees of God such as themselves existed. In other words, the angels asked the age-old question: what makes humans so special and different? Fast forward to our present age and this question is made relevant again in light of the encroaching arrival of an artificial superintelligence (ASI). Up to this point in history, humans have exceeded other creatures in various respects; now a possibility has arisen that another entity, namely ASI, will exceed humans at least on the level of intelligence and power. In relation to the age of angels, pre-modern Sunni exegesis construed human distinctiveness along the axes of reproductive knowledge and stewardship. Both brittle, distinguishing markers will disappear in the age of the ASI. Conversely, a more resilient and creative Islamic response can be derived from Ibn al-ʿArabī’s (d. 1240) views on God and the imago Dei. Inspired by the Akbarian perspective, this paper construes human distinctiveness in relation to a capacity to expansively respond to God’s love to be recognized, a response that relies on (a) imitating divine virtues that operate in counterintuitive and illogical ways, and (b) exhibiting fragility and lack rather than exceptionalism. ASI, while responding already in part to God’s love, needs to make strides towards these traits before it can answer divine love as commensurately as humans can.","venue":true},{"title":"Friendly AI will still be our master. Or, why we should not want to be the pets of super-intelligent computers","authors":["Robert Sparrow"],"published":"2023-06-13","abs_url":"https://link.springer.com/article/10.1007/s00146-023-01698-x","primary_category":"AI & SOCIETY","abstract":"When asked about humanity’s future relationship with computers, Marvin Minsky famously replied “If we’re lucky, they might decide to keep us as pets”. A number of eminent authorities continue to argue that there is a real danger that “super-intelligent” machines will enslave—perhaps even destroy—humanity. One might think that it would swiftly follow that we should abandon the pursuit of AI. Instead, most of those who purport to be concerned about the existential threat posed by AI default to worrying about what they call the “Friendly AI problem”. Roughly speaking this is the question of how we might ensure that the AI that will develop from the first AI that we create will remain sympathetic to humanity and continue to serve, or at least take account of, our interests. In this paper I draw on the “neo-republican” philosophy of Philip Pettit to argue that solving the Friendly AI problem would not change the fact that the advent of super-intelligent AI would be disastrous for humanity by virtue of rendering us the slaves of machines. A key insight of the republican tradition is that freedom requires equality of a certain sort, which is clearly lacking between pets and their owners. Benevolence is not enough. As long as AI has the power to interfere in humanity’s choices, and the capacity to do so without reference to our interests, then it will dominate us and thereby render us unfree. The pets of kind owners are still pets, which is not a status which humanity should embrace. If we really think that there is a risk that research on AI will lead to the emergence of a superintelligence, then we need to think again about the wisdom of researching AI at all.","venue":true},{"id":"2302.00843","version":7,"title":"Computational Dualism and Objective Superintelligence","authors":["Michael Timothy Bennett"],"published":"2023-02-01","updated":"2024-10-31","primary_category":"cs.AI","categories":["cs.AI","math.LO"],"abstract":"The concept of intelligent software is flawed. The behaviour of software is determined by the hardware that \"interprets\" it. This undermines claims regarding the behaviour of theorised, software superintelligence. Here we characterise this problem as \"computational dualism\", where instead of mental and physical substance, we have software and hardware. We argue that to make objective claims regarding performance we must avoid computational dualism. We propose a pancomputational alternative wherein every aspect of the environment is a relation between irreducible states. We formalise systems as behaviour (inputs and outputs), and cognition as embodied, embedded, extended and enactive. The result is cognition formalised as a part of the environment, rather than as a disembodied policy interacting with the environment through an interpreter. This allows us to make objective claims regarding intelligence, which we argue is the ability to \"generalise\", identify causes and adapt. We then establish objective upper bounds for intelligent behaviour. This suggests AGI will be safer, but more limited, than theorised.","abs_url":"https://arxiv.org/abs/2302.00843","pdf_url":"https://arxiv.org/pdf/2302.00843v7","match":"both"},{"title":"Are superintelligent robots entitled to human rights?","authors":["John-Stewart Gordon"],"published":"2022-07-10","abs_url":"https://onlinelibrary.wiley.com/doi/10.1111/rati.12346","primary_category":"Ratio 35(3): 181–193","abstract":"This paper considers relatively long‐term possibilities for the future relationship between humans and superintelligent robots (SRs). The great technological developments in fields such as artificial intelligence (AI), robotics and computer science have made it quite likely that we will see the advent of SRs towards the end of this century (or somewhat later). If SRs have a higher moral and legal status than typical adult human beings based on their greater psychological capacities, then they should also be entitled to human rights. However, even though SRs might be entitled to stronger moral and legal rights for this reason, it might nonetheless be necessary to limit their (otherwise justified) claims to avoid causing human beings to become extinct or endangered. The paper provides an argument in support of SRs' claims to human rights but also warns about the socio‐political, moral and legal implications of taking such a step.","venue":true},{"id":"2206.03487","version":3,"title":"Formalization of the principles of brain Programming (Brain Principles Programming)","authors":["E. E. Vityaev","A. G. Kolonin","A. V. Kurpatov A. A. Molchanov"],"published":"2022-05-13","updated":"2022-06-14","primary_category":"cs.AI","categories":["cs.AI"],"abstract":"In the monograph \"Strong artificial intelligence. On the Approaches to Superintelligence \" contains an overview of general artificial intelligence (AGI). As an anthropomorphic research area, it includes Brain Principles Programming (BPP) -- the formalization of universal mechanisms (principles) of the brain work with information, which are implemented at all levels of the organization of nervous tissue. This monograph contains a formalization of these principles in terms of category theory. However, this formalization is not enough to develop algorithms for working with information. In this paper, for the description and modeling of BPP, it is proposed to apply mathematical models and algorithms developed earlier, which modeling cognitive functions and base on well-known physiological, psychological and other natural science theories. The paper uses mathematical models and algorithms of the following theories: P.K.Anokhin Theory of Functional Brain Systems, Eleanor Rosch prototypical categorization theory, Bob Rehder theory of causal models and \"natural\" classification. As a result, a formalization of BPP is obtained and computer experiments demonstrating the operation of algorithms are presented.","abs_url":"https://arxiv.org/abs/2206.03487","pdf_url":"https://arxiv.org/pdf/2206.03487v3","match":"abstract"},{"title":"Will Superintelligence Lead to Spiritual Enhancement?","authors":["Ted Peters"],"published":"2022-04-26","abs_url":"https://www.mdpi.com/2077-1444/13/5/399","primary_category":"Religions 13(5): 399","abstract":"If we human beings are successful at enhancing our intelligence through technology, will this count as spiritual advance? No. Intelligence alone—whether what we are born with or what is superseded by artificial intelligence or intelligence amplification—has no built-in moral compass. Christian spirituality values love more highly than intelligence, because love orients us toward God, toward the welfare of the neighbor, and toward the common good. Spiritual advance would require orienting our enhanced intelligence toward loving God and neighbor with heart, mind (or intelligence), and soul.","venue":true},{"id":"2203.17255","version":7,"title":"A Cognitive Architecture for Machine Consciousness and Artificial Superintelligence: Thought Is Structured by the Iterative Updating of Working Memory","authors":["Jared Edward Reser"],"published":"2022-03-29","updated":"2024-11-13","primary_category":"q-bio.NC","categories":["q-bio.NC","cs.CL","cs.CV"],"abstract":"This article provides an analytical framework for how to simulate human-like thought processes within a computer. It describes how attention and memory should be structured, updated, and utilized to search for associative additions to the stream of thought. The focus is on replicating the dynamics of the mammalian working memory system, which features two forms of persistent activity: sustained firing (preserving information on the order of seconds) and synaptic potentiation (preserving information from minutes to hours). The article uses a series of figures to systematically demonstrate how the iterative updating of these working memory stores provides functional organization to behavior, cognition, and awareness. In a machine learning implementation, these two memory stores should be updated continuously and in an iterative fashion. This means each state should preserve a proportion of the coactive representations from the state before it (where each representation is an ensemble of neural network nodes). This makes each state a revised iteration of the preceding state and causes successive configurations to overlap and blend with respect to the information they contain. Thus, the set of concepts in working memory will evolve gradually and incrementally over time. Transitions between states happen as persistent activity spreads activation energy throughout the hierarchical network, searching long-term memory for the most appropriate representation to be added to the global workspace. The result is a chain of associatively linked intermediate states capable of advancing toward a solution or goal. Iterative updating is conceptualized here as an information processing strategy, a model of working memory, a theory of consciousness, and an algorithm for designing and programming artificial intelligence (AI, AGI, and ASI).","abs_url":"https://arxiv.org/abs/2203.17255","pdf_url":"https://arxiv.org/pdf/2203.17255v7","match":"title"},{"id":"2202.12710","version":3,"title":"Brain Principles Programming","authors":["Evgenii Vityaev","Anton Kolonin","Andrey Kurpatov","Artem Molchanov"],"published":"2022-02-13","updated":"2022-04-03","primary_category":"q-bio.NC","categories":["q-bio.NC","cs.AI"],"abstract":"In the monograph, STRONG ARTIFICIAL INTELLIGENCE. On the Approaches to Superintelligence, published by Sberbank, provides a cross-disciplinary review of general artificial intelligence. As an anthropomorphic direction of research, it considers Brain Principles Programming, BPP) the formalization of universal mechanisms (principles) of the brain's work with information, which are implemented at all levels of the organization of nervous tissue. This monograph provides a formalization of these principles in terms of the category theory. However, this formalization is not enough to develop algorithms for working with information. In this paper, for the description and modeling of Brain Principles Programming, it is proposed to apply mathematical models and algorithms developed by us earlier that model cognitive functions, which are based on well-known physiological, psychological and other natural science theories. The paper uses mathematical models and algorithms of the following theories: P.K.Anokhin's Theory of Functional Brain Systems, Eleonor Rosh's prototypical categorization theory, Bob Rehter's theory of causal models and natural classification. As a result, the formalization of the BPP is obtained and computer examples are given that demonstrate the algorithm's operation.","abs_url":"https://arxiv.org/abs/2202.12710","pdf_url":"https://arxiv.org/pdf/2202.12710v3","match":"abstract"},{"title":"A philosophical view on singularity and strong AI","authors":["Christian Hugo Hoffmann"],"published":"2022-01-17","abs_url":"https://link.springer.com/article/10.1007/s00146-021-01327-5","primary_category":"AI & SOCIETY","abstract":"More intellectual modesty, but also conceptual clarity is urgently needed in AI, perhaps more than in many other disciplines. AI research has been coined by hypes and hubris since its early beginnings in the 1950s. For instance, the Nobel laureate Herbert Simon predicted after his participation in the Dartmouth workshop that “machines will be capable, within 20 years, of doing any work that a man can do”. And expectations are in some circles still high to overblown today. This paper addresses the demand for conceptual clarity and introduces precise definitions of “strong AI”, “superintelligence”, the “technological singularity”, and “artificial general intelligence” which ground in the work by the computer scientist Judea Pearl and the psychologist Howard Gardner. These clarifications allow us to embed famous arguments from the philosophy of AI in a more analytic context.","venue":true},{"title":"Optimising peace through a Universal Global Peace Treaty to constrain the risk of war from a militarised artificial superintelligence","authors":["Elias G. Carayannis","John Draper"],"published":"2022-01-11","abs_url":"https://link.springer.com/article/10.1007/s00146-021-01382-y","primary_category":"AI & SOCIETY","abstract":"This article argues that an artificial superintelligence (ASI) emerging in a world where war is still normalised constitutes a catastrophic existential risk, either because the ASI might be employed by a nation–state to war for global domination, i.e., ASI-enabled warfare, or because the ASI wars on behalf of itself to establish global domination, i.e., ASI-directed warfare. Presently, few states declare war or even war on each other, in part due to the 1945 UN Charter, which states Member States should “refrain in their international relations from the threat or use of force”, while allowing for UN Security Council-endorsed military measures and self-defense. As UN Member States no longer declare war on each other, instead, only ‘international armed conflicts’ occur. However, costly interstate conflicts, both hot and cold and tantamount to wars, still take place. Further, a New Cold War between AI superpowers looms. An ASI-directed/enabled future conflict could trigger total war, including nuclear conflict, and is therefore high risk. Via conforming instrumentalism, an international relations theory, we advocate risk reduction by optimising peace through a Universal Global Peace Treaty (UGPT), contributing towards the ending of existing wars and prevention of future wars, as well as a Cyberweapons and Artificial Intelligence Convention. This strategy could influence state actors, including those developing ASIs, or an agential ASI, particularly if it values conforming instrumentalism and peace.","venue":true},{"id":"2201.02950","version":1,"title":"Arguments about Highly Reliable Agent Designs as a Useful Path to Artificial Intelligence Safety","authors":["Issa Rice","David Manheim"],"published":"2022-01-09","updated":"2022-01-09","primary_category":"cs.AI","categories":["cs.AI","cs.CY"],"abstract":"Several different approaches exist for ensuring the safety of future Transformative Artificial Intelligence (TAI) or Artificial Superintelligence (ASI) systems, and proponents of different approaches have made different and debated claims about the importance or usefulness of their work in the near term, and for future systems. Highly Reliable Agent Designs (HRAD) is one of the most controversial and ambitious approaches, championed by the Machine Intelligence Research Institute, among others, and various arguments have been made about whether and how it reduces risks from future AI systems. In order to reduce confusion in the debate about AI safety, here we build on a previous discussion by Rice which collects and presents four central arguments which are used to justify HRAD as a path towards safety of AI systems. We have titled the arguments (1) incidental utility,(2) deconfusion, (3) precise specification, and (4) prediction. Each of these makes different, partly conflicting claims about how future AI systems can be risky. We have explained the assumptions and claims based on a review of published and informal literature, along with consultation with experts who have stated positions on the topic. Finally, we have briefly outlined arguments against each approach and against the agenda overall.","abs_url":"https://arxiv.org/abs/2201.02950","pdf_url":"https://arxiv.org/pdf/2201.02950v1","match":"abstract"},{"title":"Compression, The Fermi Paradox and Artificial Super-Intelligence","authors":["Michael Timothy Bennett"],"published":"2022-01-06","abs_url":"https://link.springer.com/chapter/10.1007/978-3-030-93758-4_5","primary_category":"International Conference on Artificial General Intelligence — Artificial General Intelligence (LNCS)","abstract":"The following briefly discusses possible difficulties in communication with and control of an AGI (artificial general intelligence), building upon an explanation of The Fermi Paradox and preceding work on symbol emergence and artificial general intelligence. The latter suggests that to infer what someone means, an agent constructs a rationale for the observed behaviour of others. Communication then requires two agents labour under similar compulsions and have similar experiences (construct similar solutions to similar tasks). Any non-human intelligence may construct solutions such that any rationale for their behaviour (and thus the meaning of their signals) is outside the scope of what a human is inclined to notice or comprehend. Further, the more compressed a signal, the closer it will appear to random noise. Another intelligence may possess the ability to compress information to the extent that, to us, their signals would appear indistinguishable from noise (an explanation for The Fermi Paradox). To facilitate predictive accuracy an AGI would tend to more compressed representations of the world, making any rationale for their behaviour more difficult to comprehend for the same reason. Communication with and control of an AGI may subsequently necessitate not only human-like compulsions and experiences, but imposed cognitive impairment.","venue":true},{"id":"2112.11184","version":2,"title":"Principles for new ASI Safety Paradigms","authors":["Erland Wittkotter","Roman Yampolskiy"],"published":"2021-12-02","updated":"2022-02-14","primary_category":"cs.CY","categories":["cs.CY"],"abstract":"Artificial Superintelligence (ASI) that is invulnerable, immortal, irreplaceable, unrestricted in its powers, and above the law is likely persistently uncontrollable. The goal of ASI Safety must be to make ASI mortal, vulnerable, and law-abiding. This is accomplished by having (1) features on all devices that allow killing and eradicating ASI, (2) protect humans from being hurt, damaged, blackmailed, or unduly bribed by ASI, (3) preserving the progress made by ASI, including offering ASI to survive a Kill-ASI event within an ASI Shelter, (4) technically separating human and ASI activities so that ASI activities are easier detectable, (5) extending Rule of Law to ASI by making rule violations detectable and (6) create a stable governing system for ASI and Human relationships with reliable incentives and rewards for ASI solving humankinds problems. As a consequence, humankind could have ASI as a competing multiplet of individual ASI instances, that can be made accountable and being subjects to ASI law enforcement, respecting the rule of law, and being deterred from attacking humankind, based on humanities ability to kill-all or terminate specific ASI instances. Required for this ASI Safety is (a) an unbreakable encryption technology, that allows humans to keep secrets and protect data from ASI, and (b) watchdog (WD) technologies in which security-relevant features are being physically separated from the main CPU and OS to prevent a comingling of security and regular computation.","abs_url":"https://arxiv.org/abs/2112.11184","pdf_url":"https://arxiv.org/pdf/2112.11184v2","match":"abstract"},{"id":"2109.07899","version":1,"title":"On the Unimportance of Superintelligence","authors":["John G. Sotos"],"published":"2021-08-29","updated":"2021-08-29","primary_category":"cs.CY","categories":["cs.CY"],"abstract":"Humankind faces many existential threats, but has limited resources to mitigate them. Choosing how and when to deploy those resources is, therefore, a fateful decision. Here, I analyze the priority for allocating resources to mitigate the risk of superintelligences. Part I observes that a superintelligence unconnected to the outside world (de-efferented) carries no threat, and that any threat from a harmful superintelligence derives from the peripheral systems to which it is connected, e.g., nuclear weapons, biotechnology, etc. Because existentially-threatening peripheral systems already exist and are controlled by humans, the initial effects of a superintelligence would merely add to the existing human-derived risk. This additive risk can be quantified and, with specific assumptions, is shown to decrease with the square of the number of humans having the capability to collapse civilization. Part II proposes that biotechnology ranks high in risk among peripheral systems because, according to all indications, many humans already have the technological capability to engineer harmful microbes having pandemic spread. Progress in biomedicine and computing will proliferate this threat. ``Savant'' software that is not generally superintelligent will underpin much of this progress, thereby becoming the software responsible for the highest and most imminent existential risk -- ahead of hypothetical risk from superintelligences. The analysis concludes that resources should be preferentially applied to mitigating the risk of peripheral systems and savant software. Concerns about superintelligence are at most secondary, and possibly superfluous.","abs_url":"https://arxiv.org/abs/2109.07899","pdf_url":"https://arxiv.org/pdf/2109.07899v1","match":"both"},{"title":"Feeding the Beast: Superintelligence, Corporate Capitalism and the End of Humanity","authors":["Dominic Leggett"],"published":"2021-07-21","abs_url":"https://doi.org/10.1145/3461702.3462581","primary_category":"Proceedings of the 2021 AAAI/ACM Conference on AI, Ethics, and Society (AIES 2021), pp. 727–735","abstract":null,"venue":true},{"title":"Society 5.0: A Japanese Concept for a Superintelligent Society","authors":["Carolina Narvaez Rojas","Gustavo Adolfo Alomia Peñafiel","Diego Fernando Loaiza Buitrago","Carlos Andrés Tavera Romero"],"published":"2021-06-09","abs_url":"https://www.mdpi.com/2071-1050/13/12/6567","primary_category":"Sustainability 13(12): 6567","abstract":"This document discusses the Japanese context of Society 5.0. Based on a society-centered approach, Society 5.0 seeks to take advantage of technological advances to finally solve the problems that currently threaten Japan, such as aging, birth rates and lack of competitiveness, among others. Additionally, another objective is to contribute to the progress of the country and develop the foundations for a better world, in which no individual can be excluded from the technological advances of our current society, to achieve this goal, the Sustainable Development Goals (SDG) have been developed. SDGs seek to assess the methods of use of modern technology and thus find the best strategies and tools to use it in a way that guarantees sustainability within the framework of a new society that demands constant renovations.","venue":true},{"title":"What overarching ethical principle should a superintelligent AI follow?","authors":["Atle Ottesen Søvik"],"published":"2021-05-27","abs_url":"https://link.springer.com/article/10.1007/s00146-021-01229-6","primary_category":"AI & SOCIETY","abstract":"What is the best overarching ethical principle to give a possible future superintelligent machine, given that we do not know what the best ethics are today or in the future? Eliezer Yudkowsky has suggested that a superintelligent AI should have as its goal to carry out the coherent extrapolated volition of humanity (CEV), the most coherent way of combining human goals. The article discusses some problems with this proposal and some alternatives suggested by Nick Bostrom. A slightly different proposal is then suggested, which I argue solves the problems better than Yudkowsky’s proposal.","venue":true},{"title":"Ethical Regulators and Super-Ethical Systems","authors":["Mick Ashby"],"published":"2020-12-09","abs_url":"https://www.mdpi.com/2079-8954/8/4/53","primary_category":"Systems 8(4): 53","abstract":"This paper combines the good regulator theorem with the law of requisite variety and seven other requisites that are necessary and sufficient for a cybernetic regulator to be effective and ethical. The ethical regulator theorem provides a basis for systematically evaluating and improving the adequacy of existing or proposed designs for systems that make decisions that can have ethical consequences; regardless of whether the regulators are humans, machines, cyberanthropic hybrids, organizations, or government institutions. The theorem is used to define an ethical design process that has potentially far‐reaching implications for society. A six‐level framework is proposed for classifying cybernetic and superintelligent systems, which highlights the existence of a possibility‐space bifurcation in our future time‐line. The implementation of “super‐ethical” systems is identified as an urgent imperative for humanity to avoid the danger that superintelligent machines might lead to a technological dystopia. It is proposed to define third‐order cybernetics as the cybernetics of ethical systems. Concrete actions, a grand challenge, and a vision of a super‐ethical society are proposed to help steer the future of the human race and our wonderful planet towards a realistically achievable minimum viable cyberanthropic utopia.","venue":true},{"title":"Artificial superintelligence and its limits: why AlphaZero cannot become a general agent","authors":["Karim Jebari","Joakim Lundborg"],"published":"2020-10-13","abs_url":"https://link.springer.com/article/10.1007/s00146-020-01070-3","primary_category":"AI & SOCIETY","abstract":"An intelligent machine surpassing human intelligence across a wide set of skills has been proposed as a possible existential catastrophe (i.e., an event comparable in value to that of human extinction). Among those concerned about existential risk related to artificial intelligence (AI), it is common to assume that AI will not only be very intelligent, but also be a general agent (i.e., an agent capable of action in many different contexts). This article explores the characteristics of machine agency, and what it would mean for a machine to become a general agent. In particular, it does so by articulating some important differences between belief and desire in the context of machine agency. One such difference is that while an agent can by itself acquire new beliefs through learning, desires need to be derived from preexisting desires or acquired with the help of an external influence. Such influence could be a human programmer or natural selection. We argue that to become a general agent, a machine needs productive desires, or desires that can direct behavior across multiple contexts. However, productive desires cannot sui generis be derived from non-productive desires. Thus, even though general agency in AI could in principle be created by human agents, general agency cannot be spontaneously produced by a non-general AI agent through an endogenous process (i.e. self-improvement). In conclusion, we argue that a common AI scenario, where general agency suddenly emerges in a non-general agent AI, such as DeepMind’s superintelligent board game AI AlphaZero, is not plausible.","venue":true},{"title":"Public Policy and Superintelligent AI: A Vector Field Approach","authors":["Nick Bostrom","Allan Dafoe","Carrick Flynn"],"published":"2020-09-17","abs_url":"https://doi.org/10.1093/oso/9780190905033.003.0011","primary_category":"Ethics of Artificial Intelligence, ed. S. Matthew Liao (Oxford University Press)","abstract":"We consider the speculative prospect of superintelligent AI and its normative implications for governance and global policy. Machine superintelligence would be a transformative development that would present a host of political challenges and opportunities. This paper identifies a set of distinctive features of this hypothetical policy context, from which we derive a correlative set of policy desiderata—considerations that should be given extra weight in long-term AI policy compared to in other policy contexts. Our contribution describes a desiderata “vector field” showing the directional change from a variety of possible normative baselines or policy positions. The focus on directional normative change should make our findings relevant to a wide range of actors, although the development of concrete policy options that meet these abstractly formulated desiderata will require further work.","venue":true},{"title":"IN ALGORITHMS WE TRUST: MAGICAL THINKING, SUPERINTELLIGENT AI AND QUANTUM COMPUTING","authors":["Nathan Schradle"],"published":"2020-09-02","abs_url":"https://www.zygonjournal.org/article/id/14681/","primary_category":"Zygon: Journal of Religion and Science 55(3): 733–747","abstract":"This article analyzes current attitudes toward artificial intelligence (AI) and quantum computing and argues that they represent a modern‐day form of magical thinking. It proposes that AI and quantum computing are thus excellent examples of the ways that traditional distinctions between religion, science, and magic fail to account for the vibrancy and energy that surround modern technologies.","venue":true},{"id":"2008.04071","version":1,"title":"On Controllability of AI","authors":["Roman V. Yampolskiy"],"published":"2020-07-18","updated":"2020-07-18","primary_category":"cs.CY","categories":["cs.CY","cs.AI"],"abstract":"Invention of artificial general intelligence is predicted to cause a shift in the trajectory of human civilization. In order to reap the benefits and avoid pitfalls of such powerful technology it is important to be able to control it. However, possibility of controlling artificial general intelligence and its more advanced version, superintelligence, has not been formally established. In this paper, we present arguments as well as supporting evidence from multiple domains indicating that advanced AI can't be fully controlled. Consequences of uncontrollability of AI are discussed with respect to future of humanity and research on AI, and AI safety and security.","abs_url":"https://arxiv.org/abs/2008.04071","pdf_url":"https://arxiv.org/pdf/2008.04071v1","match":"abstract"},{"id":"2007.03616","version":1,"title":"Artificial Stupidity","authors":["Michael Falk"],"published":"2020-07-01","updated":"2020-07-01","primary_category":"cs.CY","categories":["cs.CY","cs.AI"],"abstract":"Public debate about AI is dominated by Frankenstein Syndrome, the fear that AI will become superhuman and escape human control. Although superintelligence is certainly a possibility, the interest it excites can distract the public from a more imminent concern: the rise of Artificial Stupidity (AS). This article discusses the roots of Frankenstein Syndrome in Mary Shelley's famous novel of 1818. It then provides a philosophical framework for analysing the stupidity of artificial agents, demonstrating that modern intelligent systems can be seen to suffer from 'stupidity of judgement'. Finally it identifies an alternative literary tradition that exposes the perils and benefits of AS. In the writings of Edmund Spenser, Jonathan Swift and E.T.A. Hoffmann, ASs replace, oppress or seduce their human users. More optimistically, Joseph Furphy and Laurence Sterne imagine ASs that can serve human intellect as maps or as pipes. These writers provide a strong counternarrative to the myths that currently drive the AI debate. They identify ways in which even stupid artificial agents can evade human control, for instance by appealing to stereotypes or distancing us from reality. And they underscore the continuing importance of the literary imagination in an increasingly automated society.","abs_url":"https://arxiv.org/abs/2007.03616","pdf_url":"https://arxiv.org/pdf/2007.03616v1","match":"abstract"},{"id":"1909.12152","version":1,"title":"Superintelligence Safety: A Requirements Engineering Perspective","authors":["Hermann Kaindl","Jonas Ferdigg"],"published":"2019-09-26","updated":"2019-09-26","primary_category":"cs.AI","categories":["cs.AI","cs.SE"],"abstract":"Under the headline \"AI safety\", a wide-reaching issue is being discussed, whether in the future some \"superhuman artificial intelligence\" / \" superintelligence \" could could pose a threat to humanity. In addition, the late Steven Hawking warned that the rise of robots may be disastrous for mankind. A major concern is that even benevolent superhuman artificial intelligence (AI) may become seriously harmful if its given goals are not exactly aligned with ours, or if we cannot specify precisely its objective function. Metaphorically, this is compared to king Midas in Greek mythology, who expressed the wish that everything he touched should turn to gold, but obviously this wish was not specified precisely enough. In our view, this sounds like requirements problems and the challenge of their precise formulation. (To our best knowledge, this has not been pointed out yet.) As usual in requirements engineering (RE), ambiguity or incompleteness may cause problems. In addition, the overall issue calls for a major RE endeavor, figuring out the wishes and the needs with regard to a superintelligence, which will in our opinion most likely be a very complex software-intensive system based on AI. This may even entail theoretically defining an extended requirements problem.","abs_url":"https://arxiv.org/abs/1909.12152","pdf_url":"https://arxiv.org/pdf/1909.12152v1","match":"both"},{"id":"1908.01766","version":1,"title":"Seeding the Singularity for A.I","authors":["Pavel Kraikivski"],"published":"2019-08-04","updated":"2019-08-04","primary_category":"cs.AI","categories":["cs.AI"],"abstract":"The singularity refers to an idea that once a machine having an artificial intelligence surpassing the human intelligence capacity is created, it will trigger explosive technological and intelligence growth. I propose to test the hypothesis that machine intelligence capacity can grow autonomously starting with an intelligence comparable to that of bacteria - microbial intelligence. The goal will be to demonstrate that rapid growth in intelligence capacity can be realized at all in artificial computing systems. I propose the following three properties that may allow an artificial intelligence to exhibit a steady growth in its intelligence capacity: (i) learning with the ability to modify itself when exposed to more data, (ii) acquiring new functionalities (skills), and (iii) expanding or replicating itself. The algorithms must demonstrate a rapid growth in skills of dataprocessing and analysis and gain qualitatively different functionalities, at least until the current computing technology supports their scalable development. The existing algorithms that already encompass some of these or similar properties, as well as missing abilities that must yet be implemented, will be reviewed in this work. Future computational tests could support or oppose the hypothesis that artificial intelligence can potentially grow to the level of superintelligence which overcomes the limitations in hardware by producing necessary processing resources or by changing the physical realization of computation from using chip circuits to using quantum computing principles.","abs_url":"https://arxiv.org/abs/1908.01766","pdf_url":"https://arxiv.org/pdf/1908.01766v1","match":"abstract"},{"title":"Future-Ready Strategic Oversight of Multiple Artificial Superintelligence-Enabled Adaptive Learning Systems via Human-Centric Explainable AI-Empowered Predictive Optimizations of Educational Outcomes","authors":["Meng-Leong How"],"published":"2019-07-31","abs_url":"https://www.mdpi.com/2504-2289/3/3/46","primary_category":"Big Data and Cognitive Computing 3(3): 46","abstract":"Artificial intelligence-enabled adaptive learning systems (AI-ALS) have been increasingly utilized in education. Schools are usually afforded the freedom to deploy the AI-ALS that they prefer. However, even before artificial intelligence autonomously develops into artificial superintelligence in the future, it would be remiss to entirely leave the students to the AI-ALS without any independent oversight of the potential issues. For example, if the students score well in formative assessments within the AI-ALS but subsequently perform badly in paper-based posttests, or if the relentless algorithm of a particular AI-ALS is suspected of causing undue stress for the students, they should be addressed by educational stakeholders. Policy makers and educational stakeholders should collaborate to analyze the data from multiple AI-ALS deployed in different schools to achieve strategic oversight. The current paper provides exemplars to illustrate how this future-ready strategic oversight could be implemented using an artificial intelligence-based Bayesian network software to analyze the data from five dissimilar AI-ALS, each deployed in a different school. Besides using descriptive analytics to reveal potential issues experienced by students within each AI-ALS, this human-centric AI-empowered approach also enables explainable predictive analytics of the students’ learning outcomes in paper-based summative assessments after training is completed in each AI-ALS.","venue":true},{"title":"The Future Impact of Artificial Intelligence on Humans and Human Rights","authors":["Steven Livingston","Mathias Risse"],"published":"2019-06-07","abs_url":"https://www.cambridge.org/core/journals/ethics-and-international-affairs/article/abs/future-impact-of-artificial-intelligence-on-humans-and-human-rights/2016EDC9A61F68615EBF9AFA8DE91BF8","primary_category":"Ethics & International Affairs","abstract":"What are the implications of artificial intelligence (AI) on human rights in the next three decades? Precise answers to this question are made difficult by the rapid rate of innovation in AI research and by the effects of human practices on the adaption of new technologies. Precise answers are also challenged by imprecise usages of the term “AI.” There are several types of research that all fall under this general term. We begin by clarifying what we mean by AI. Most of our attention is then focused on the implications of artificial general intelligence (AGI), which entail that an algorithm or group of algorithms will achieve something like superintelligence. While acknowledging that the feasibility of superintelligence is contested, we consider the moral and ethical implications of such a potential development. What do machines owe humans and what do humans owe superintelligent machines?","venue":true},{"id":"1905.04288","version":1,"title":"Growth, degrowth, and the challenge of artificial superintelligence","authors":["Salvador Pueyo"],"published":"2019-05-03","updated":"2019-05-03","primary_category":"cs.CY","categories":["cs.CY"],"abstract":"The implications of technological innovation for sustainability are becoming increasingly complex with information technology moving machines from being mere tools for production or objects of consumption to playing a role in economic decision making. This emerging role will acquire overwhelming importance if, as a growing body of literature suggests, artificial intelligence is underway to outperform human intelligence in most of its dimensions, thus becoming \" superintelligence \". Hitherto, the risks posed by this technology have been framed as a technical rather than a political challenge. With the help of a thought experiment, this paper explores the environmental and social implications of superintelligence emerging in an economy shaped by neoliberal policies. It is argued that such policies exacerbate the risk of extremely adverse impacts. The experiment also serves to highlight some serious flaws in the pursuit of economic efficiency and growth per se, and suggests that the challenge of superintelligence cannot be separated from the other major environmental and social challenges, demanding a fundamental transformation along the lines of degrowth. Crucially, with machines outperforming them in their functions, there is little reason to expect economic elites to be exempt from the threats that superintelligence would pose in a neoliberal context, which opens a door to overcoming vested interests that stand in the way of social change toward sustainability and equity.","abs_url":"https://arxiv.org/abs/1905.04288","pdf_url":"https://arxiv.org/pdf/1905.04288v1","match":"both"},{"title":"Risk management standards and the active management of malicious intent in artificial superintelligence","authors":["Patrick Bradley"],"published":"2019-04-11","abs_url":"https://link.springer.com/article/10.1007/s00146-019-00890-2","primary_category":"AI & SOCIETY","abstract":"The likely near future creation of artificial superintelligence carries significant risks to humanity. These risks are difficult to conceptualise and quantify, but malicious use of existing artificial intelligence by criminals and state actors is already occurring and poses risks to digital security, physical security and integrity of political systems. These risks will increase as artificial intelligence moves closer to superintelligence. While there is little research on risk management tools used in artificial intelligence development, the current global standard for risk management, ISO 31000:2018, is likely used extensively by developers of artificial intelligence technologies. This paper argues that risk management has a common set of vulnerabilities when applied to artificial superintelligence which cannot be resolved within the existing framework and alternative approaches must be developed. Some vulnerabilities are similar to issues posed by malicious threat actors such as professional criminals and terrorists. Like these malicious actors, artificial superintelligence will be capable of rendering mitigation ineffective by working against countermeasures or attacking in ways not anticipated by the risk management process. Criminal threat management recognises this vulnerability and seeks to guide and block the intent of malicious threat actors as an alternative to risk management. An artificial intelligence treachery threat model that acknowledges the failings of risk management and leverages the concepts of criminal threat management and artificial stupidity is proposed. This model identifies emergent malicious behaviour and allows intervention against negative outcomes at the moment of artificial intelligence’s greatest vulnerability.","venue":true},{"title":"Global Solutions vs. Local Solutions for the AI Safety Problem","authors":["Alexey Turchin","David Denkenberger","Brian Patrick Green"],"published":"2019-02-20","abs_url":"https://www.mdpi.com/2504-2289/3/1/16","primary_category":"Big Data and Cognitive Computing 3(1): 16","abstract":"There are two types of artificial general intelligence (AGI) safety solutions: global and local. Most previously suggested solutions are local: they explain how to align or “box” a specific AI (Artificial Intelligence), but do not explain how to prevent the creation of dangerous AI in other places. Global solutions are those that ensure any AI on Earth is not dangerous. The number of suggested global solutions is much smaller than the number of proposed local solutions. Global solutions can be divided into four groups: 1. No AI: AGI technology is banned or its use is otherwise prevented; 2. One AI: the first superintelligent AI is used to prevent the creation of any others; 3. Net of AIs as AI police: a balance is created between many AIs, so they evolve as a net and can prevent any rogue AI from taking over the world; 4. Humans inside AI: humans are augmented or part of AI. We explore many ideas, both old and new, regarding global solutions for AI safety. They include changing the number of AI teams, different forms of “AI Nanny” (non-self-improving global control AI system able to prevent creation of dangerous AIs), selling AI safety solutions, and sending messages to future AI. Not every local solution scales to a global solution or does it ethically and safely. The choice of the best local solution should include understanding of the ways in which it will be scaled up. Human-AI teams or a superintelligent AI Service as suggested by Drexler may be examples of such ethically scalable local solutions, but the final choice depends on some unknown variables such as the speed of AI progress.","venue":true},{"title":"Reframing Superintelligence: Comprehensive AI Services as General Intelligence","authors":["K. Eric Drexler"],"published":"2019","abs_url":"https://www.fhi.ox.ac.uk/wp-content/uploads/Reframing_Superintelligence_FHI-TR-2019-1.1-1.pdf","primary_category":"Future of Humanity Institute Technical Report #2019-1","abstract":null,"venue":true},{"id":"1811.03009","version":1,"title":"Uploading Brain into Computer: Whom to Upload First?","authors":["Yana B. Feygin","Kelly Morris","Roman V. Yampolskiy"],"published":"2018-10-27","updated":"2018-10-27","primary_category":"cs.CY","categories":["cs.CY","cs.AI"],"abstract":"The final goal of the intelligence augmentation process is a complete merger of biological brains and computers allowing for integration and mutual enhancement between computer's speed and memory and human's intelligence. This process, known as uploading, analyzes human brain in detail sufficient to understand its working patterns and makes it possible to simulate said brain on a computer. As it is likely that such simulations would quickly evolve or be modified to achieve superintelligence it is very important to make sure that the first brain chosen for such a procedure is a suitable one. In this paper, we attempt to answer the question: Whom to upload first?","abs_url":"https://arxiv.org/abs/1811.03009","pdf_url":"https://arxiv.org/pdf/1811.03009v1","match":"abstract"},{"title":"Countering Superintelligence Misinformation","authors":["Seth D. Baum"],"published":"2018-09-30","abs_url":"https://www.mdpi.com/2078-2489/9/10/244","primary_category":"Information 9(10): 244","abstract":"Superintelligence is a potential type of future artificial intelligence (AI) that is significantly more intelligent than humans in all major respects. If built, superintelligence could be a transformative event, with potential consequences that are massively beneficial or catastrophic. Meanwhile, the prospect of superintelligence is the subject of major ongoing debate, which includes a significant amount of misinformation. Superintelligence misinformation is potentially dangerous, ultimately leading bad decisions by the would-be developers of superintelligence and those who influence them. This paper surveys strategies to counter superintelligence misinformation. Two types of strategies are examined: strategies to prevent the spread of superintelligence misinformation and strategies to correct it after it has spread. In general, misinformation can be difficult to correct, suggesting a high value of strategies to prevent it. This paper is the first extended study of superintelligence misinformation. It draws heavily on the study of misinformation in psychology, political science, and related fields, especially misinformation about global warming. The strategies proposed can be applied to lay public attention to superintelligence, AI education programs, and efforts to build expert consensus.","venue":true},{"title":"Moral Status of Digital Agents: Acting Under Uncertainty","authors":["Abhishek Mishra"],"published":"2018-08-29","abs_url":"https://link.springer.com/chapter/10.1007/978-3-319-96448-5_30","primary_category":"Philosophy and Theory of Artificial Intelligence 2017 (PT-AI 2017), Studies in Applied Philosophy, Epistemology and Rational Ethics","abstract":"This paper addresses how to act towards digital agents while uncertain about their moral status. It focuses specifically on the problem of how to act towards simulated minds operated by an artificial superintelligence (ASI). This problem can be treated as a sub-set of the larger problems of AI-safety (how to ensure a desirable outcome after the emergence of ASI) and also invokes debates about the grounds of moral status. The paper presents a formal structure for solving the problem by first constraining it as a sub-problem to the AI-safety problem, and then suggesting a decision-theoretic approach to how this problem can be solved under uncertainty about what the true grounds of moral status are, and whether such simulations do possess these relevant grounds. The paper ends by briefly suggesting a way to generalize the approach.","venue":true},{"title":"Friendly Superintelligent AI: All You Need Is Love","authors":["Michael Prinzing"],"published":"2018-08-29","abs_url":"https://link.springer.com/chapter/10.1007/978-3-319-96448-5_31","primary_category":"Philosophy and Theory of Artificial Intelligence 2017 (PT-AI 2017), Studies in Applied Philosophy, Epistemology and Rational Ethics","abstract":"There is a non-trivial chance that sometime in the (perhaps somewhat distant) future, someone will build an artificial general intelligence that will surpass human-level cognitive proficiency and go on to become “superintelligent”, vastly outperforming humans. The advent of superintelligent AI has great potential, for good or ill. It is therefore imperative that we find a way to ensure— long before one arrives—that any superintelligence we build will consistently act in ways congenial to our interests. This is a very difficult challenge in part because most of the final goals we could give an AI admit of so-called “perverse instantiations”. I propose a novel solution to this puzzle: instruct the AI to love humanity. The proposal is compared with Yudkowsky’s Coherent Extrapolated Volition, and Bostrom’s Moral Modeling proposals.","venue":true},{"title":"Superintelligence Skepticism as a Political Tool","authors":["Seth D. Baum"],"published":"2018-08-22","abs_url":"https://www.mdpi.com/2078-2489/9/9/209","primary_category":"Information 9(9): 209","abstract":"This paper explores the potential for skepticism about artificial superintelligence to be used as a tool for political ends. Superintelligence is AI that is much smarter than humans. Superintelligence does not currently exist, but it has been proposed that it could someday be built, with massive and potentially catastrophic consequences. There is substantial skepticism about superintelligence, including whether it will be built, whether it would be catastrophic, and whether it is worth current attention. To date, superintelligence skepticism appears to be mostly honest intellectual debate, though some of it may be politicized. This paper finds substantial potential for superintelligence skepticism to be (further) politicized, due mainly to the potential for major corporations to have a strong profit motive to downplay concerns about superintelligence and avoid government regulation. Furthermore, politicized superintelligence skepticism is likely to be quite successful, due to several factors including the inherent uncertainty of the topic and the abundance of skeptics. The paper’s analysis is based on characteristics of superintelligence and the broader AI sector, as well as the history and ongoing practice of politicized skepticism on other science and technology issues, including tobacco, global warming, and industrial chemicals. The paper contributes to literatures on politicized skepticism and superintelligence governance.","venue":true},{"title":"Who Owns My Autonomous Vehicle? Ethics and Responsibility in Artificial and Human Intelligence","authors":["JOHN HARRIS"],"published":"2018-08-06","abs_url":"https://www.cambridge.org/core/journals/cambridge-quarterly-of-healthcare-ethics/article/abs/who-owns-my-autonomous-vehicle-ethics-and-responsibility-in-artificial-and-human-intelligence/0DFD45A804F21686A997D0C4D3C47558","primary_category":"Cambridge Quarterly of Healthcare Ethics","abstract":"This article investigates both the claims made for, and the dangers or opportunities posed by, the development of (allegedly), aspiring or “would-be” autonomous vehicles and other artificially superintelligent machines. It also examines the dilemmas posed by the fact that these individuals might develop ideas above their station. These ideas may also limit or challenge the legitimacy of the proposed management and safety strategies that might be devised to limit the ways in which they might function or malfunction.","venue":true},{"title":"Hybrid Strategies Towards Safe “Self-Aware” Superintelligent Systems","authors":["Nadisha-Marie Aliman","Leon Kester"],"published":"2018-07-21","abs_url":"https://link.springer.com/chapter/10.1007/978-3-319-97676-1_1","primary_category":"International Conference on Artificial General Intelligence — Artificial General Intelligence (LNCS)","abstract":"Against the backdrop of increasing progresses in AI research paired with a rise of AI applications in decision-making processes, security-critical domains as well as in ethically relevant frames, a large-scale debate on possible safety measures encompassing corresponding long-term and short-term issues has emerged across different disciplines. One pertinent topic in this context which has been addressed by various AI Safety researchers is e.g. the AI alignment problem for which no final consensus has been achieved yet. In this paper, we present a multidisciplinary toolkit of AI Safety strategies combining considerations from AI and Systems Engineering as well as from Cognitive Science with a security mindset as often relevant in Cybersecurity. We elaborate on how AGI “Self-awareness” could complement different AI Safety measures in a framework extended by a jointly performed Human Enhancement procedure. Our analysis suggests that this hybrid framework could contribute to undertake the AI alignment problem from a new holistic perspective through security-building synergetic effects emerging thereof and could help to increase the odds of a possible safe future transition towards superintelligent systems.","venue":true},{"title":"Will There Be Superintelligence and Would It Hate Us?","authors":["Yorick Wilks"],"published":"2017-12-28","abs_url":"https://ojs.aaai.org/aimagazine/index.php/aimagazine/article/view/2726","primary_category":"AI Magazine 38(4): 65–70","abstract":"Bostrom’s Superintelligence (SI) is a wide-ranging essay (2016) that has raised important questions about the future of intelligent machines and the possible malign developments they may undergo. But, and perhaps surprisingly, it is not about technical developments in artificial intelligence (AI) nor a philosophical analysis of the concept of SI. There is little of either of these in it, which is largely an extended and stimulating essay on economics, decision theory and other forms of social science, all held together by the unsubstantiated hypothesis of “superintelligence” that belongs more to science fiction than AI. AI may well in some future produce undesirable social effects — the Internet itself could already be such a development — but there is as yet no reason to think they could be on the massive and end-of-civilization scale Bostrom so confidently predicts.","venue":true},{"title":"Modeling and Interpreting Expert Disagreement About Artificial Superintelligence","authors":["Seth D Baum","Anthony M Barrett","Roman V Yampolskiy"],"published":"2017-12-27","abs_url":"https://www.informatica.si/index.php/informatica/article/view/1812","primary_category":"Informatica 41(4), Special Issue on Superintelligence","abstract":"Artificial superintelligence (ASI) is artificial intelligence (AI) with capabilities that are significantly greater than human capabilities across a wide range of domains. A hallmark of the ASI issue is disagreement among experts. This paper demonstrates and discusses methodological options for modeling and interpreting expert disagreement about the risk of ASI catastrophe. Using a new model called ASI-PATH, the paper models a well-documented recent disagreement between Nick Bostrom and Ben Goertzel, two distinguished ASI experts. Three points of disagreement are considered: (1) the potential for humans to evaluate the values held by an AI, (2) the potential for humans to create an AI with values that humans would consider desirable, and (3) the potential for an AI to create for itself values that humans would consider desirable. An initial quantitative analysis shows that accounting for variation in expert judgment can have a large effect on estimates of the risk of ASI catastrophe. The risk estimates can in turn inform ASI risk management strategies, which the paper demonstrates via an analysis of the strategy of AI confinement. The paper find the optimal strength of AI confinement to depend on the balance of risk parameters (1) and (2).","venue":true},{"title":"Conceptual-Linguistic Superintelligence","authors":["David J. Jilk"],"published":"2017-12-27","abs_url":"https://www.informatica.si/index.php/informatica/article/view/1875","primary_category":"Informatica 41(4), Special Issue on Superintelligence","abstract":"We argue that artificial intelligence capable of sustaining an uncontrolled intelligence explosion must have a conceptual-linguistic faculty with substantial functional similarity to the human faculty. We then argue for three subsidiary claims: first, that detecting the presence of such a faculty will be an important indicator of imminent superintelligence; second, that such a superintelligence will, in creating further increases in intelligence, both face and consider the same sorts of existential risks that humans face today; third, that such a superintelligence is likely to assess and question its own values, purposes, and drives.","venue":true},{"title":"Superintelligence As a Cause or Cure For Risks of Astronomical Suffering","authors":["Kaj Sotala","Lukas Gloor"],"published":"2017-12-27","abs_url":"https://www.informatica.si/index.php/informatica/article/view/1877","primary_category":"Informatica 41(4), Special Issue on Superintelligence","abstract":"Discussions about the possible consequences of creating superintelligence have included the possibility of existential risk , often understood mainly as the risk of human extinction. We argue that suffering risks (s-risks) , where an adverse outcome would bring about severe suffering on an astronomical scale, are risks of a comparable severity and probability as risks of extinction. Preventing them is the common interest of many different value systems. Furthermore, we argue that in the same way as superintelligent AI both contributes to existential risk but can also help prevent it, superintelligent AI can both be a suffering risk or help avoid it. Some types of work aimed at making superintelligent AI safe will also help prevent suffering risks, and there may also be a class of safeguards for AI that helps specifically against s-risks.","venue":true},{"title":"Artificial Intelligence in Life Extension: from Deep Learning to Superintelligence","authors":["Mikhail Batin","Alexey Turchin","Markov Sergey","Alisa Zhila","David Denkenberger"],"published":"2017-12-27","abs_url":"https://www.informatica.si/index.php/informatica/article/view/1797","primary_category":"Informatica 41(4), Special Issue on Superintelligence","abstract":"In this paper we focus on the most efficacious AI applications for life extension and anti-aging at three expected stages of AI development: narrow AI, AGI and superintelligence. First, we overview the existing research and commercial work performed by a select number of startups and academic projects. We find that at the current stage of “narrow” AI, the most promising areas for life extension are geroprotector-combination discovery, detection of aging biomarkers, and personalized anti-aging therapy. These advances could help currently living people reach longevity escape velocity and survive until more advanced AI appears. When AI comes close to human level, the main contribution to life extension will come from AI integration with humans through brain-computer interfaces, integrated AI assistants capable of autonomously diagnosing and treating health issues, and cyber systems embedded into human bodies. Lastly, we speculate about the more remote future, when AI reaches the level of superintelligence and such life-extension methods as uploading human minds and creating nanotechnological bodies may become possible, thus lowering the probability of human death close to zero. We conclude that medical AI based superintelligence is intrinsically safer than, say, military AI, as it may help humans to evolve into part of the future superintelligence via brain augmentation, uploading, and a network of self-improving humans. Medical AI’s value system is focused on human benefit.","venue":true},{"title":"How feasible is the rapid development of artificial superintelligence?","authors":["Kaj Sotala"],"published":"2017-10-24","abs_url":"https://iopscience.iop.org/article/10.1088/1402-4896/aa90e8","primary_category":"Physica Scripta","abstract":"What kinds of fundamental limits are there in how capable artificial intelligence (AI) systems might become? Two questions in particular are of interest: (1) How much more capable could AI become relative to humans, and (2) how easily could superhuman capability be acquired? To answer these questions, we will consider the literature on human expertise and intelligence, discuss its relevance for AI, and consider how AI could improve on humans in two major aspects of thought and expertise, namely simulation and pattern recognition. We find that although there are very real limits to prediction, it seems like AI could still substantially improve on human intelligence. Export citation and abstract BibTeX RIS Next article in issue","venue":true},{"title":"The problem of superintelligence: political, not technological","authors":["Wolfhart Totschnig"],"published":"2017-08-09","abs_url":"https://link.springer.com/article/10.1007/s00146-017-0753-0","primary_category":"AI & SOCIETY","abstract":"The thinkers who have reflected on the problem of a coming superintelligence have generally seen the issue as a technological problem, a problem of how to control what the superintelligence will do. I argue that this approach is probably mistaken because it is based on questionable assumptions about the behavior of intelligent agents and, moreover, potentially counterproductive because it might, in the end, bring about the existential catastrophe that it is meant to prevent. I contend that the problem posed by a future superintelligence will likely be a political problem, that is, one of establishing a peaceful form of coexistence with other intelligent agents in a situation of mutual vulnerability, and not a technological problem of control.","venue":true},{"title":"Superintelligent AI and the Postbiological Cosmos Approach1","authors":["Susan Schneider"],"published":"2017-07-13","abs_url":"https://www.cambridge.org/core/books/abs/what-is-life-on-earth-and-beyond/superintelligent-ai-and-the-postbiological-cosmos-approach1/DD01B1E8BCADB580F2BACE4E4D21F65D","primary_category":"What is Life? On Earth and Beyond (Cambridge University Press)","abstract":null,"venue":true},{"title":"Superintelligent AI and Skepticism","authors":["Joseph Corabi"],"published":"2017-06-01","abs_url":"https://jeet.ieet.org/index.php/home/article/view/63","primary_category":"Journal of Ethics and Emerging Technologies","abstract":"It has become fashionable to worry about the development of superintelligent AI that results in the destruction of humanity. This worry is not without merit, but it may be overstated. This paper explores some previously undiscussed reasons to be optimistic that, even if superintelligent AI does arise, it will not destroy us. These have to do with the possibility that a superintelligent AI will become mired in skeptical worries that its superintelligence cannot help it to solve. I argue that superintelligent AIs may lack the psychological idiosyncracies that allow humans to act in the face of skeptical problems, and so as a result they may become paralyzed in the face of these problems in a way that humans are not.","venue":true},{"title":"Games between humans and AIs","authors":["Stephen J. DeCanio"],"published":"2017-05-31","abs_url":"https://link.springer.com/article/10.1007/s00146-017-0732-5","primary_category":"AI & SOCIETY","abstract":"Various potential strategic interactions between a “strong” Artificial intelligence (AI) and humans are analyzed using simple 2 × 2 order games, drawing on the New Periodic Table of those games developed by Robinson and Goforth (The topology of the 2 × 2 games: a new periodic table. Routledge, London, 2005 ). Strong risk aversion on the part of the human player(s) leads to shutting down the AI research program, but alternative preference orderings by the human and the AI result in Nash equilibria with interesting properties. Some of the AI-Human games have multiple equilibria, and in other cases Pareto-improvement over the Nash equilibrium could be attained if the AI’s behavior towards humans could be guaranteed to be benign. The preferences of a superintelligent AI cannot be known in advance, but speculation is possible as to its ranking of alternative states of the world, and how it might assimilate the accumulated wisdom (and folly) of humanity.","venue":true},{"id":"1702.08495","version":2,"title":"Don't Fear the Reaper: Refuting Bostrom's Superintelligence Argument","authors":["Sebastian Benthall"],"published":"2017-02-27","updated":"2017-03-04","primary_category":"cs.AI","categories":["cs.AI"],"abstract":"In recent years prominent intellectuals have raised ethical concerns about the consequences of artificial intelligence. One concern is that an autonomous agent might modify itself to become \" superintelligent \" and, in supremely effective pursuit of poorly specified goals, destroy all of humanity. This paper considers and rejects the possibility of this outcome. We argue that this scenario depends on an agent's ability to rapidly improve its ability to predict its environment through self-modification. Using a Bayesian model of a reasoning agent, we show that there are important limitations to how an agent may improve its predictive ability through self-modification alone. We conclude that concern about this artificial intelligence outcome is misplaced and better directed at policy questions around data access and storage.","abs_url":"https://arxiv.org/abs/1702.08495","pdf_url":"https://arxiv.org/pdf/1702.08495v2","match":"title"},{"id":"1702.08529","version":1,"title":"Multi-agent systems and decentralized artificial superintelligence","authors":["S. Ponomarev","A. E. Voronkov"],"published":"2017-02-27","updated":"2017-02-27","primary_category":"cs.MA","categories":["cs.MA"],"abstract":"Multi-agents systems communication is a technology, which provides a way for multiple interacting intelligent agents to communicate with each other and with environment. Multiple-agent systems are used to solve problems that are difficult for solving by individual agent. Multiple-agent communication technologies can be used for management and organization of computing fog and act as a global, distributed operating system. In present publication we suggest technology, which combines decentralized P2P BOINC general-purpose computing tasks distribution, multiple-agents communication protocol and smart-contract based rewards, powered by Ethereum blockchain. Such system can be used as distributed P2P computing power market, protected from any central authority. Such decentralized market can further be updated to system, which learns the most efficient way for software-hardware combinations usage and optimization. Once system learns to optimize software-hardware efficiency it can be updated to general-purpose distributed intelligence, which acts as combination of single-purpose AI.","abs_url":"https://arxiv.org/abs/1702.08529","pdf_url":"https://arxiv.org/pdf/1702.08529v1","match":"title"},{"title":"Singularitarianism and schizophrenia","authors":["Vassilis Galanos"],"published":"2016-10-31","abs_url":"https://link.springer.com/article/10.1007/s00146-016-0679-y","primary_category":"AI & SOCIETY","abstract":"Given the contemporary ambivalent standpoints toward the future of artificial intelligence, recently denoted as the phenomenon of Singularitarianism, Gregory Bateson’s core theories of ecology of mind, schismogenesis, and double bind, are hereby revisited, taken out of their respective sociological, anthropological, and psychotherapeutic contexts and recontextualized in the field of Roboethics as to a twofold aim: (a) the proposal of a rigid ethical standpoint toward both artificial and non-artificial agents, and (b) an explanatory analysis of the reasons bringing about such a polarized outcome of contradictory views in regard to the future of robots. Firstly, the paper applies the Batesonian ecology of mind for constructing a unified roboethical framework which endorses a flat ontology embracing multiple forms of agency, borrowing elements from Floridi’s information ethics, classic virtue ethics, Felix Guattari’s ecosophy, Braidotti’s posthumanism, and the Japanese animist doctrine of Rinri. The proposed framework wishes to act as a pragmatic solution to the endless dispute regarding the nature of consciousness or the natural/artificial dichotomy and as a further argumentation against the recognition of future artificial agency as a potential existential threat. Secondly, schismogenic analysis is employed to describe the emergence of the hostile human–robot cultural contact, tracing its origins in the early scientific discourse of man–machine symbiosis up to the contemporary countermeasures against superintelligent agents. Thirdly, Bateson’s double bind theory is utilized as an analytic methodological tool of humanity’s collective agency, leading to the hypothesis of collective schizophrenic symptomatology, due to the constancy and intensity of confronting messages emitted by either proponents or opponents of artificial intelligence. The double bind’s treatment is the mirroring “therapeutic double bind,” and the article concludes in proposing the conceptual pragmatic imperative necessary for such a condition to follow: humanity’s conscience of habitualizing danger and familiarization with its possible future extinction, as the result of a progressive blurrification between natural and artificial agency, succeeded by a totally non-organic intelligent form of agency.","venue":true},{"id":"1609.02009","version":1,"title":"Non-Evolutionary Superintelligences Do Nothing, Eventually","authors":["Telmo Menezes"],"published":"2016-09-07","updated":"2016-09-07","primary_category":"cs.AI","categories":["cs.AI","cs.CY"],"abstract":"There is overwhelming evidence that human intelligence is a product of Darwinian evolution. Investigating the consequences of self-modification, and more precisely, the consequences of utility function self-modification, leads to the stronger claim that not only human, but any form of intelligence is ultimately only possible within evolutionary processes. Human-designed artificial intelligences can only remain stable until they discover how to manipulate their own utility function. By definition, a human designer cannot prevent a superhuman intelligence from modifying itself, even if protection mechanisms against this action are put in place. Without evolutionary pressure, sufficiently advanced artificial intelligences become inert by simplifying their own utility function. Within evolutionary processes, the implicit utility function is always reducible to persistence, and the control of superhuman intelligences embedded in evolutionary processes is not possible. Mechanisms against utility function self-modification are ultimately futile. Instead, scientific effort toward the mitigation of existential risks from the development of superintelligences should be in two directions: understanding consciousness, and the complex dynamics of evolutionary systems.","abs_url":"https://arxiv.org/abs/1609.02009","pdf_url":"https://arxiv.org/pdf/1609.02009v1","match":"both"},{"id":"1609.00331","version":3,"title":"Verifier Theory and Unverifiability","authors":["Roman V. Yampolskiy"],"published":"2016-09-01","updated":"2016-10-25","primary_category":"cs.AI","categories":["cs.AI","cs.CR","cs.SE"],"abstract":"Despite significant developments in Proof Theory, surprisingly little attention has been devoted to the concept of proof verifier. In particular, the mathematical community may be interested in studying different types of proof verifiers (people, programs, oracles, communities, superintelligences) as mathematical objects. Such an effort could reveal their properties, their powers and limitations (particularly in human mathematicians), minimum and maximum complexity, as well as self-verification and self-reference issues. We propose an initial classification system for verifiers and provide some rudimentary analysis of solved and open problems in this important domain. Our main contribution is a formal introduction of the notion of unverifiability, for which the paper could serve as a general citation in domains of theorem proving, as well as software and AI verification.","abs_url":"https://arxiv.org/abs/1609.00331","pdf_url":"https://arxiv.org/pdf/1609.00331v3","match":"abstract"},{"title":"Outer space and cyber space: meeting ET in the cloud","authors":["Ted Peters"],"published":"2016-08-02","abs_url":"https://www.cambridge.org/core/journals/international-journal-of-astrobiology/article/abs/outer-space-and-cyber-space-meeting-et-in-the-cloud/8ECD4760E38594EC28138617840F1689","primary_category":"International Journal of Astrobiology","abstract":"What justifies the astrobiologist's search for post-biological or machine-intelligence in outer space? Four assumptions borrowed from transhumanism (H+) seem to be at work: (1) it is reasonable to speculate that life on Earth will evolve in the direction of post-biological intelligence; (2) if extraterrestrials have evolved longer than we on Earth, then they will be more scientifically and technologically advanced; (3) superintelligence, computer uploads of brains, and dis-embodied mind belong together; and (4) evolutionary progress is guided by the drive toward increased intelligence. When subjected to critical review, these assumptions prove to be weak. Most importantly, evolutionary biologists do not support the idea that evolution is internally directed toward increased intelligence. Without this assumption, justifying the search for ET more intelligent than earthlings is anaemic. Nevertheless, one can still hope that in the near future we will be communicating with new neighbours in the Milky Way. Can sheer hope inspire science?","venue":true},{"id":"1607.07730","version":1,"title":"A Model of Pathways to Artificial Superintelligence Catastrophe for Risk and Decision Analysis","authors":["Anthony M. Barrett","Seth D. Baum"],"published":"2016-07-25","updated":"2016-07-25","primary_category":"cs.AI","categories":["cs.AI"],"abstract":"An artificial superintelligence (ASI) is artificial intelligence that is significantly more intelligent than humans in all respects. While ASI does not currently exist, some scholars propose that it could be created sometime in the future, and furthermore that its creation could cause a severe global catastrophe, possibly even resulting in human extinction. Given the high stakes, it is important to analyze ASI risk and factor the risk into decisions related to ASI research and development. This paper presents a graphical model of major pathways to ASI catastrophe, focusing on ASI created via recursive self-improvement. The model uses the established risk and decision analysis modeling paradigms of fault trees and influence diagrams in order to depict combinations of events and conditions that could lead to AI catastrophe, as well as intervention options that could decrease risks. The events and conditions include select aspects of the ASI itself as well as the human process of ASI research, development, and management. Model structure is derived from published literature on ASI risk. The model offers a foundation for rigorous quantitative evaluation and decision making on the long-term risk of ASI catastrophe.","abs_url":"https://arxiv.org/abs/1607.07730","pdf_url":"https://arxiv.org/pdf/1607.07730v1","match":"both"},{"id":"1607.00913","version":1,"title":"Superintelligence cannot be contained: Lessons from Computability Theory","authors":["Manuel Alfonseca","Manuel Cebrian","Antonio Fernandez Anta","Lorenzo Coviello","Andres Abeliuk","Iyad Rahwan"],"published":"2016-07-04","updated":"2016-07-04","primary_category":"cs.CY","categories":["cs.CY","cs.AI"],"abstract":"Superintelligence is a hypothetical agent that possesses intelligence far surpassing that of the brightest and most gifted human minds. In light of recent advances in machine intelligence, a number of scientists, philosophers and technologists have revived the discussion about the potential catastrophic risks entailed by such an entity. In this article, we trace the origins and development of the neo-fear of superintelligence, and some of the major proposals for its containment. We argue that such containment is, in principle, impossible, due to fundamental limits inherent to computing itself. Assuming that a superintelligence will contain a program that includes all the programs that can be executed by a universal Turing machine on input potentially as complex as the state of the world, strict containment requires simulations of such a program, something theoretically (and practically) infeasible.","abs_url":"https://arxiv.org/abs/1607.00913","pdf_url":"https://arxiv.org/pdf/1607.00913v1","match":"both"},{"title":"Don’t Worry about Superintelligence","authors":["Nicholas Agar"],"published":"2016-02-01","abs_url":"https://jeet.ieet.org/index.php/home/article/view/52","primary_category":"Journal of Ethics and Emerging Technologies","abstract":"This paper responds to Nick Bostrom’s suggestion that the threat of a human-unfriendly superintelligence should lead us to delay or rethink progress in AI. I allow that progress in AI presents problems that we are currently unable to solve. However, we should distinguish between currently unsolved problems for which there are rational expectations of solutions and currently unsolved problems for which no such expectation is appropriate. The problem of a human-unfriendly superintelligence belongs to the first category. It is rational to proceed on that assumption that we will solve it. These observations do not reduce to zero the existential threat from superintelligence. But we should not permit fear of very improbable negative outcomes to delay the arrival of the expected benefits from AI.","venue":true},{"title":"Superintelligence: Fears, Promises and Potentials","authors":["Ben Goertzel"],"published":"2015-12-01","abs_url":"https://jeet.ieet.org/index.php/home/article/view/48","primary_category":"Journal of Ethics and Emerging Technologies","abstract":"Oxford philosopher Nick Bostrom, in his recent and celebrated book Superintelligence , argues that advanced AI poses a potentially major existential risk to humanity, and that advanced AI development should be heavily regulated and perhaps even restricted to a small set of government-approved researchers. Bostrom’s ideas and arguments are reviewed and explored in detail, and compared with the thinking of three other current thinkers on the nature and implications of AI: Eliezer Yudkowsky of the Machine Intelligence Research Institute (formerly Singularity Institute for AI), and David Weinbaum (Weaver) and Viktoras Veitas of the Global Brain Institute. Relevant portions of Yudkowsky’s book Rationality: From AI to Zombies are briefly reviewed, and it is found that nearly all the core ideas of Bostrom’s work appeared previously or concurrently in Yudkowsky’s thinking. However, Yudkowsky often presents these shared ideas in a more plain-spoken and extreme form, making clearer the essence of what is being claimed. For instance, the elitist strain of thinking that one sees in the background in Bostrom is plainly and openly articulated in Yudkowsky, with many of the same practical conclusions (e.g. that it may well be best if advanced AI is developed in secret by a small elite group). Bostrom and Yudkowsky view intelligent systems through the lens of reinforcement learning – they view them as “reward-maximizers” and worry about what happens when a very powerful and intelligent reward-maximizer is paired with a goal system that gives rewards for achieving foolish goals like tiling the universe with paperclips. Weinbaum and Veitas’s recent paper “Open-Ended Intelligence” presents a starkly alternative perspective on intelligence, viewing it as centered not on reward maximization, but rather on complex self-organization and self-transcending development that occurs in close coupling with a complex environment that is also ongoingly self-organizing, in only partially knowable ways. It is concluded that Bostrom and Yudkowsky’s arguments for existential risk have some logical foundation, but are often presented in an exaggerated way. For instance, formal arguments whose implication is that the “worst case scenarios” for advanced AI development are extremely dire, are often informally discussed as if they demonstrated the likelihood, rather than just the possibility, of highly negative outcomes. And potential dangers of reward-maximizing AI are taken as problems with AI in general, rather than just as problems of the reward-maximization paradigm as an approach to building superintelligence. If one views past, current, and future intelligence as “open-ended,” in the vernacular of Weaver and Veitas, the potential dangers no longer appear to loom so large, and one sees a future that is wide-open, complex and uncertain, just as it has always been.","venue":true},{"title":"Why AI Doomsayers are Like Sceptical Theists and Why it Matters","authors":["John Danaher"],"published":"2015-04-25","abs_url":"https://link.springer.com/article/10.1007/s11023-015-9365-y","primary_category":"Minds and Machines","abstract":"An advanced artificial intelligence (a “superintelligence”) could pose a significant existential risk to humanity. Several research institutes have been set-up to address those risks. And there is an increasing number of academic publications analysing and evaluating their seriousness. Nick Bostrom’s superintelligence: paths, dangers, strategies represents the apotheosis of this trend. In this article, I argue that in defending the credibility of AI risk, Bostrom makes an epistemic move that is analogous to one made by so-called sceptical theists in the debate about the existence of God. And while this analogy is interesting in its own right, what is more interesting are its potential implications. It has been repeatedly argued that sceptical theism has devastating effects on our beliefs and practices. Could it be that AI-doomsaying has similar effects? I argue that it could. Specifically, and somewhat paradoxically, I argue that it could amount to either a reductio of the doomsayers position, or an important and additional reason to join their cause. I use this paradox to suggest that the modal standards for argument in the superintelligence debate need to be addressed.","venue":true},{"title":"Superintelligence: Paths, Dangers, Strategies","authors":["Nick Bostrom"],"published":"2014-09-03","abs_url":"https://global.oup.com/academic/product/superintelligence-9780199678112","primary_category":"Oxford University Press","abstract":null,"venue":true},{"id":"1405.3378","version":1,"title":"The \"crisis of noosphere\" as a limiting factor to achieve the point of technological singularity","authors":["Rafael Lahoz-Beltra"],"published":"2014-05-14","updated":"2014-05-14","primary_category":"cs.CY","categories":["cs.CY"],"abstract":"One of the most significant developments in the history of human being is the invention of a way of keeping records of human knowledge, thoughts and ideas. In 1926, the work of several thinkers such as Edouard Le Roy, Vladimir Vernadsky and Teilhard de Chardin led to the concept of noosphere, thus the idea that human cognition and knowledge transforms the biosphere coming to be something like the planet's thinking layer. At present, is commonly accepted by some thinkers that the Internet is the medium that brings life to noosphere. According to Vinge and Kurzweil's technological singularity hypothesis, noosphere would be in the future the natural environment in which 'human-machine superintelligence ' emerges after to reach the point of technological singularity. In this paper we show by means of a numerical model the impossibility that our civilization reaches the point of technological singularity in the near future. We propose that this point may be reached when Internet data centers are based on \"computer machines\" to be more effective in terms of power consumption than current ones. We speculate about what we have called 'Nooscomputer' or N-computer a hypothetical machine which would consume far less power allowing our civilization to reach the point of technological singularity.","abs_url":"https://arxiv.org/abs/1405.3378","pdf_url":"https://arxiv.org/pdf/1405.3378v1","match":"abstract"},{"title":"Risks and Mitigation Strategies for Oracle AI","authors":["Stuart Armstrong"],"published":"2013","abs_url":"https://link.springer.com/chapter/10.1007/978-3-642-31674-6_25","primary_category":"Philosophy and Theory of Artificial Intelligence (PT-AI 2011), Studies in Applied Philosophy, Epistemology and Rational Ethics","abstract":"There is no strong reason to believe human level intelligence represents an upper limit of the capacity of artificial intelligence, should it be realized. This poses serious safety issues, since a superintelligent system would have great power to direct the future according to its possibly flawed goals or motivation systems. Oracle AIs (OAI), confined AIs that can only answer questions, are one particular approach to this problem. However even Oracles are not particularly safe: humans are still vulnerable to traps, social engineering, or simply becoming dependent on the OAI. But OAIs are still strictly safer than general AIs, and there are many extra layers of precautions we can add on top of these. This paper looks at some of them and analyses their strengths and weaknesses.","venue":true},{"title":"What to Do with the Singularity Paradox?","authors":["Roman V. Yampolskiy"],"published":"2013","abs_url":"https://link.springer.com/chapter/10.1007/978-3-642-31674-6_30","primary_category":"Philosophy and Theory of Artificial Intelligence (PT-AI 2011), Studies in Applied Philosophy, Epistemology and Rational Ethics","abstract":"The paper begins with an introduction of the Singularity Paradox, an observation that: “Superintelligent machines are feared to be too dumb to possess commonsense”. Ideas from leading researchers in the fields of philosophy, mathematics, economics, computer science and robotics regarding the ways to address said paradox are reviewed and evaluated. Suggestions are made regarding the best way to handle the Singularity Paradox.","venue":true},{"title":"The Superintelligent Will: Motivation and Instrumental Rationality in Advanced Artificial Agents","authors":["Nick Bostrom"],"published":"2012-06-13","abs_url":"https://link.springer.com/article/10.1007/s11023-012-9281-3","primary_category":"Minds and Machines","abstract":"This paper discusses the relation between intelligence and motivation in artificial agents, developing and briefly arguing for two theses. The first, the orthogonality thesis , holds (with some caveats) that intelligence and final goals (purposes) are orthogonal axes along which possible artificial intellects can freely vary—more or less any level of intelligence could be combined with more or less any final goal. The second, the instrumental convergence thesis , holds that as long as they possess a sufficient level of intelligence, agents having any of a wide range of final goals will pursue similar intermediary goals because they have instrumental reasons to do so. In combination, the two theses help us understand the possible range of behavior of superintelligent agents, and they point to some potential dangers in building such an agent.","venue":true},{"title":"Thinking Inside the Box: Controlling and Using an Oracle AI","authors":["Stuart Armstrong","Anders Sandberg","Nick Bostrom"],"published":"2012-06-06","abs_url":"https://link.springer.com/article/10.1007/s11023-012-9282-2","primary_category":"Minds and Machines","abstract":"There is no strong reason to believe that human-level intelligence represents an upper limit of the capacity of artificial intelligence, should it be realized. This poses serious safety issues, since a superintelligent system would have great power to direct the future according to its possibly flawed motivation system. Solving this issue in general has proven to be considerably harder than expected. This paper looks at one particular approach, Oracle AI. An Oracle AI is an AI that does not act in the world except by answering questions. Even this narrow approach presents considerable challenges. In this paper, we analyse and critique various methods of controlling the AI. In general an Oracle AI might be safer than unrestricted AI, but still remains potentially dangerous.","venue":true},{"title":"A Computability Argument Against Superintelligence","authors":["Jiří Wiedermann"],"published":"2012-02-18","abs_url":"https://link.springer.com/article/10.1007/s12559-012-9124-9","primary_category":"Cognitive Computation","abstract":"Using the contemporary view of computing exemplified by recent models and results from non-uniform complexity theory, we investigate the computational power of cognitive systems. We show that in accordance with the so-called extended Turing machine paradigm such systems can be modelled as non-uniform evolving interactive systems whose computational power surpasses that of the classical Turing machines. Our results show that there is an infinite hierarchy of cognitive systems. Within this hierarchy, there are systems achieving and surpassing the human intelligence level. Any intelligence level surpassing the human intelligence is called the superintelligence level. We will argue that, formally, from a computation viewpoint the human-level intelligence is upper-bounded by the \\(\\Upsigma_2\\) class of the Arithmetical Hierarchy. In this class, there are problems whose complexity grows faster than any computable function and, therefore, not even exponential growth of computational power can help in solving such problems, or reach the level of superintelligence.","venue":true}]}