Pivocajs

85 karmaJoined Dec 2017

Message

Bio

Vojta Kovarik. AI alignment and game theory researcher.

Posts
1

Sorted by New

In extremely high-stakes scenarios, it's ok not to maximise expected utility

Pivocajs

· 2y ago · 2m read

Comments
26

Topic contributions
2

Experts' AI timelines are longer than you have been told?

Pivocajs14d3

I want to flag that even with short timelines and selfish goals, the terms of the bet seem like a bad deal.

If, until the end of 2028, Metaculus' question about superintelligent AI:
Resolves non-ambiguously, I transfer to you 10 k January-2025-$ in the month after that in which the question resolved.
Does not resolve, you transfer to me 10 k January-2025-$ in January 2029. As before, I plan to donate my profits to animal welfare organisations.

Reason: Many people with short timelines also tend to put high probability on superintelligent AI being bad news (eg, me). From that point if view, an over-simplified interpretation of the terms is:

Either we get SAI by 2028, in which case I am dead (and get 10k).
Or we don't, in which case I have to pay 10k.

If you wanted to account for this, the bet should be modified somehow. EG, you give me 10k now, and if [the question didn't resolve] / [I am alive] by January 2029, I send you your 10k back and pay you 10k * x -- where your proposal corresponds to x=1. (FWIW, I personally wouldn't take the bet for x=1. But I would start thinking about it for x=0.5 or so.)

Cooperative AI: Three things that confused me as a beginner (and my current understanding)

Pivocajs3mo1

What (if any) is the overlap of cooperative AI, AI ethics, and AI safety? Perhaps preventing catastrophic harm that is somehow tied to failures of fairness or inclusion?

I imagine that failures as Moloch / runaway capitalism / you get what you can measure would qualify. (Or more precisely, harms caused by these would include things that AI Ethics is concerned about, in a way that Cooperative AI / AI Safety also tries to prevent.)

Should you work at a frontier AI company?

Pivocajs5mo1

I think the summary at the start of this post is too easy to misinterpret as "if you think of yourself as a smart and moral person, it's ok to go for these companies".

(None of the things the summary says seem false. But the overall impression seems too vulnerable to rationalisation along the lines of "surely I would not fall prey to these bad incentives". When reality is probably that most people fall prety to them. So at the minimum, it might be more fair to change the recommendation to something like "it's complicated, but err on the side of not joining" or "it's complicated, but we wouldn't recommend this for 95% of people who can get a job at these companies"^[1].

^{^}
Or whatever qualifier you think is fair. The main point is to make it clear that the warnings apply to the reader as well, not just to "all the other people".

Matthew_Barnett's Quick takes

Pivocajs1y1

In my opinion, the main relevant alternative to this view is to be partial to the human species, as opposed to being partial to either one's current generation, or oneself. And I think the human species is kind of a weird category to be partial to, relative to those other things. Do you disagree?

I agree with this.

the best way to advance your own values is generally to actually "be there" when AI happens.

I (strongly) disagree with this. Me being alive is a relatively small part of my values. And since I am not the director of the world, me personally being around to influence things is unlikely to have a decisive impact on things I value.

In more detail: Sure, all else being equal, me being there when AI happens is mildly helpful. But the outcome of building AI seems to be a function of, among other things, (i) values of the people building it + (ii) how much reflection they can do on those values + (iii) the environment dynamics these people are subject to (e.g., the current race dynamics between AI companies). And over time, I expect the potential decrease in (i) to be far outweighed by gains in (ii) and (iii).

The first issue is about (i), that it is not actually me building the AGI, either now or in the future. But I am willing to grant that (all else being equal) current generation is more likely to have values closer to my values.
However, I expect that the factors are (ii) and (iii) are just as influential. Regarding (ii), it seems we keep making progress at philosophy, ethics, etc, and to me, this currently far outweighs the value drift in (i).
Regarding (iii), my impression is that the current situation is so bad that it can't get much worse, and we might as well wait. This of course depends on how likely you think we are likely to get a bad outcome if we either (a) get superintelligence without additional progress on alignment or (b) get widespread human-level AI with no progress on alignment, institution design, etc.

Matthew_Barnett's Quick takes

Pivocajs1y3

My personal reason for not digging into this is that my naive model of how good the AI future is: quality_of_future * amount_of_the_stuff. And there is distinction I haven't seen you acknowledged: while high "quality" doesn't require humans to be around, I ultimately judge quality by my values. (Thing being conscious is an example. But this also includes things like not copy-pasting the same thing all over, not wiping out aliens, and presumably many other things I am not aware of. IIRC Yudkowsky talks about cosmopolitanism being a human value.) Because of this, my impression is that if we hand over the future to a random AI, the "quality" will be very low. And so we can currently have a much larger impact by focusing on increasing the quality. Which we can do by delaying "handing over the future to AI" and picking a good AI to hand over to. IE, alignment.

(Still, I agree it would be nice if there was a better analysis of this, which exposed the assumptions.)

TED talk on Moloch and AI

Pivocajs1y2

In terms of feedback/reaction: I work on AI alignment, game theory, and cooperative AI, so Moloch is basically my key concern. And from that position, I highly approve of the overall talk, and of all of the content in particular --- except for one point, where I felt a bit so-so. And that is the part about what the company leaders can do to help the situation.

The key thing is 9:58-10:09 ("We need leaders who are willing to flip the Moloch's playbook. ...") , but I think this part then changes how people interpret 10:59-10:11 ("Perhaps companies can start competing over who ... "). I don't mean to say that I strongly disagree here --- rather, I mean that this part seems objectively speculative, which was in contrast with everything else in the talk (which seemed super solid).

More specifically, the talk's formulation suggested to me that the key thing is whether the leaders would be willing to not play the Moloch game. In contrast, it seems quite possible that this by itself wouldn't help at all, for example because they would just get fired if they tried. My personal guess is that "the key thing" is affordance the leaders have for not playing the Moloch game / the costs they incur for doing so. Or perhaps the combination of this and the willingness to not play the Moloch game. And this is also how I would frame the 10:59-10:11 part --- that we should try to make it such that the companies can compete on those other things that turn this into a race to the top. (As opposed to "the companies should compete on those other things".)

Downsides of Small Organizations in EA

Pivocajs2y3

Re “Middle management is toxic, we should avoid it.”:

I want to flag that: your counterargument here does not properly address the points from Middle Manager Hell / the Immoral Mazes sequences. (Less constructively, "Middle management being toxic" seems like a quite weak version of the arguments against large orgs. Which suggests that your counterargument might not work against the stronger version. More constructively, one difference between current EA structure and large orgs is that small EA orgs are not married to a single funder. This imo reduces the "toxicity" you might otherwise get by the invectives structure in large companies. There might be other important differences; I just haven't thought about this enough.)

All that said, perhaps we can get the best of the both worlds by using larger orgs for some things but not all? And inventing some tools that make it easier to get the benefits you want without all of the costs? (Example: something that allows people to temporarily/tentatively switch jobs without having to deal with all the paperwork.)

Don't Interpret Prediction Market Prices as Probabilities

Pivocajs2y1

Just to highlight a particular example: suppose you have a prediction market on "How much will be inflation of USD over the next 2 years?", that is priced in USD.

Why "just make an agent which cares only about binary rewards" doesn't work.

Pivocajs2y3

I suggest editing the post by adding a tl;dr section to the top of the post. Or maybe change the title to something like Why "just make an agent which cares only about binary rewards" doesn't work.

Reasoning: To me, the considerations in the post mostly read as rehashing standard arguments, which one should be familiar if they thought about the problem themselves, or went through AGI Safety Fundamentals, etc. It might be interesting to some people, but it would be good to have the clear indication that this isn't novel.

Also: When I read the start of the post, I went "obviously this doesn't work". Then I spent several minutes reading the post to see where the flaw in your argument is, and point it out. Only to find that your conclusion is "yeah, this doesn't help". If you edit the post, you might save other people from wasting their time in a similar manner :-).

If your AGI x-risk estimates are low, what scenarios make up the bulk of your expectations for an OK outcome?

Answer by PivocajsApr 27, 20233

I am at high P(doom|AGI pre-2035), but not at near-certainty. Say, 75% but not 99.9%.

The reason for that is that I find both "fast takeoff takeover" and "continous multipolar takeoff" scenarios plausible (with no decisive evidence for one or the other). In "continuous multipolar takeoff", you still get superintelligences running around. However, they would be "superintelligent with respect civilization-2023" but not necessarily wrt civilization-then. And for the standard somewhat-well-thought-out AI takeover arguments to aply, you need to be superintelligent wrt civilization-then.

Two disclaimers: (1) Just because you don't get discontinuity in influence around human level does not mean you can't get it later. In my book, world can look "Christiano-like", until suddenly it looks "Yudkowsky-like". (2) Even if we never get AI singleton, things can still go horribly wrong (ie, Christiano's what failure looks like). But imo those scenarios are much harder to reason about, and we have haven't thought them out in enough detail to justify high certainty of either outcome.

My intuitive aggregation of this gives, say, 80% P(doom this century|AGI pre-2035). On top of that, I add some 5-10% on "I am so wrong about some of this that even the high-level reasoning doesn't apply". (Which includes being wrong about where the burden of proofs, and priors, lie for P(doom|AGI).) And that puts me at the (ass-) number 75%.

Pivocajs

Bio

Posts 1

Comments26

Topic contributions2

Posts
1

Comments
26

Topic contributions
2