Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

> Air gap the computer and control the inputs and outputs. We can formally prove what a system is capable of. That fixes the problem.

Debates about superhuman AI have focused quite a lot on what it would mean to "control the inputs and outputs" while still being able to get some kind of benefit from the AI.

You can indeed formally prove that a computer will or won't do certain things, so you could use that for isolation purposes. But in order to be useful, the AI needs to interact with people and/or the world in some way. Otherwise it might as well be switched off or never have been built in the first place.

If it's really a superhuman intelligence with superhuman knowledge about the world, then interacting with people is where the risk creeps back in, because the AI could make suggestions, recommendations, requests, offers, promises, or threats. Although there are plenty of ideas about limiting the nature of questions and answers, having some kind of separate person or machine judge whether information from the AI's communications should be used or how, or limiting what the AI is programmed to attempt to do or how, none of these measures are straightforward to formally prove correct in the way that simpler isolation properties are.

If we made contact with intelligent aliens, would formal proofs of correctness of the computers through which we (say, exclusively) communicate with them guarantee that they couldn't massively disrupt our society by means of what they had to say?



Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: