Best response
In game theory, the best response is the strategy (or strategies) which produces the most favorable outcome for a player, taking other players' strategies as given (Fudenberg & Tirole 1991, p. 29; Gibbons 1992, pp. 33–49). The concept of a best response is central to John Nash's best-known contribution, the Nash equilibrium, the point at which each player in a game has selected the best response (or one of the best responses) to the other players' strategies (Nash 1950).
Correspondence
Reaction correspondences, also known as best response correspondences, are used in the proof of the existence of mixed strategy Nash equilibria (Fudenberg & Tirole 1991, Section 1.3.B; Osborne & Rubinstein 1994, Section 2.2). Reaction correspondences are not "reaction functions" since functions must only have one value per argument, and many reaction correspondences will be undefined, i.e., a vertical line, for some opponent strategy choice. One constructs a correspondence , for each player from the set of opponent strategy profiles into the set of the player's strategies. So, for any given set of opponent's strategies , represents player 's best responses to .
2x2 정상 폼 게임 모두에 대한 대응 대응은 유닛 스퀘어 전략 공간에서 각 플레이어의 라인으로 그릴 수 있다. 그림 1~3 그래프는 스태그헌트 게임에 대한 최상의 대응이다. 그림 1의 점선은 플레이어 X가 스태그(x축에 표시)를 플레이할 확률의 함수로서 플레이어 Y가 '스태그'(Y축에 표시)를 플레이할 최적의 확률을 보여준다. 그림 2에서 점선은 플레이어 Y가 스태그(Y축에 표시)를 플레이할 확률의 함수로서 플레이어 X가 '스택'(X축에 표시)을 플레이할 최적의 확률을 보여준다. 그림 2는 두 플레이어의 최고 반응이 그림 3에서 일치하는 지점에서 나시 평형도를 표시하기 위해 이전 그래프에 겹쳐질 수 있도록 반대 축에 있는 독립형 및 반응 변수를 표시한다.
대칭 2x2 게임에는 세 가지 독특한 반응 대응 형태가 있는데, 세 가지 유형의 대칭 2x2 게임 각각에 한 가지씩 있다: 조정 게임, 디스코딩 게임, 지배적인 전략을 가진 게임(두 동작에 대해 항상 보상이 동일한 사소한 네 번째 경우는 실제로 게임 이론적인 문제가 아니다). 모든 대칭 2x2 게임은 이 세 가지 형태 중 하나를 취하게 될 것이다.
코디네이션 게임
남녀의 스태그헌트, 배틀 등 두 선수가 같은 전략을 선택할 때 가장 높은 점수를 받는 게임을 코디네이션 게임이라고 한다. 이 게임들은 그림 3과 같은 모양의 반작용 서신을 가지고 있는데, 그림 3은 왼쪽 아래 구석에 나시 평형 1개, 오른쪽 위쪽에 나시 평형 1개, 그리고 다른 두 개 사이의 대각선을 따라 어딘가에 혼합된 나시 평형이 있다.
반조정 게임
플레이어가 반대 전략을 선택할 때 가장 높은 점수를 받는 치킨 게임, 매독 게임과 같은 게임, 즉 디스코르드(discoordinate)를 반조정 게임이라고 한다. 이들은 조정 게임과 반대 방향으로 교차하는 리액션 대응(그림 4)을 가지고 있는데, 좌우 상단 모서리에 각각 하나씩 있는 3개의 나시 평형(Nash 평형)을 가지고 있는데, 한 명의 플레이어가 하나의 전략을 선택하고 다른 플레이어는 반대 전략을 선택한다. 세 번째 나시 평형은 왼쪽 아래에서 오른쪽 상단 모서리에 대각선을 따라 놓여 있는 혼합 전략이다. 그 중 어느 것이 어느 것인지 선수들이 모를 경우 플레이가 왼쪽 아래부터 오른쪽 위 대각선까지 제한돼 있어 혼합형 내시는 진화적으로 안정된 전략(ESS)이다. 그렇지 않으면 상관없는 비대칭이 존재한다고 하며, 코너인 내쉬 평형도는 ESS이다.
Games with dominated strategies
Games with dominated strategies have reaction correspondences which only cross at one point, which will be in either the bottom left, or top right corner in payoff symmetric 2x2 games. For instance, in the single-play prisoner's dilemma, the "Cooperate" move is not optimal for any probability of opponent Cooperation. Figure 5 shows the reaction correspondence for such a game, where the dimensions are "Probability play Cooperate", the Nash equilibrium is in the lower left corner where neither player plays Cooperate. If the dimensions were defined as "Probability play Defect", then both players best response curves would be 1 for all opponent strategy probabilities and the reaction correspondences would cross (and form a Nash equilibrium) at the top right corner.
Other (payoff asymmetric) games
A wider range of reaction correspondences shapes is possible in 2x2 games with payoff asymmetries. For each player there are five possible best response shapes, shown in Figure 6. From left to right these are: dominated strategy (always play 2), dominated strategy (always play 1), rising (play strategy 2 if probability that the other player plays 2 is above threshold), falling (play strategy 1 if probability that the other player plays 2 is above threshold), and indifferent (both strategies play equally well under all conditions).
대칭 2x2 게임의 가능한 유형은 4가지에 불과하지만(이 중 1게임은 사소한 게임임), 플레이어당 5가지 최상의 응답 곡선은 더 많은 수의 비대칭 게임 유형을 허용한다. 이것들 중 많은 것들이 서로 진정으로 다르지 않다. 치수는 논리적으로 동일한 대칭적인 게임을 생산하기 위해 다시 정의할 수 있다(전략 1과 2의 교환명).
매칭 페니
보상의 비대칭성을 가진 잘 알려진 게임 중 하나는 페니 게임과 매칭 페니 게임이다. 이 게임에서는 y차원에 그래프로 표시된 한 행 플레이어가 조정(머리 또는 꼬리 둘 다 선택)하면 승리하고 x축에 표시된 다른 행 플레이어가 디스코르디스트하면 승리한다. Y선수의 대응 서신은 조정 게임의 대응 서신이고, X선수의 대응 서신은 디스코딩 게임의 대응 서신이다. 유일한 내시 평형은 양쪽 선수가 각각 0.5개의 확률로 머리와 꼬리를 독자적으로 선택하는 혼합 전략의 조합이다.
역학
진화 게임 이론에서 최고의 대응 역학은 전략 업데이트 규칙의 한 부류를 나타내며, 다음 라운드에서 플레이어 전략은 모집단의 일부 하위 집합에 대한 최상의 반응에 의해 결정된다. 일부 예는 다음과 같다.
- 대규모 모집단 모델에서 플레이어는 어떤 전략이 모집단 전체에 대한 최상의 반응인지에 따라 확률적으로 다음 행동을 선택한다.
- 공간적 모델에서 플레이어는 모든 이웃에게 가장 잘 반응하는 액션(다음 라운드에서)을 선택한다(Ellison 1993).
중요한 것은, 이러한 모델에서 플레이어는 다음 라운드에서 가장 높은 수익을 얻을 수 있는 최고의 반응만을 선택한다는 것이다. 선수들은 다음 라운드에서 전략을 선택하는 것이 경기에서의 미래 플레이에 미칠 영향을 고려하지 않는다. 이 제약조건은 종종 근시안적인 최선의 반응이라고 불리는 역동적인 규칙을 초래한다.
In the theory of potential games, best response dynamics refers to a way of finding a Nash equilibrium by computing the best response for every player:
Theorem: In any finite potential game, best response dynamics always converge to a Nash equilibrium. (Nisan et al. 2007, Section 19.3.2)
Smoothed
Instead of best response correspondences, some models use smoothed best response functions. These functions are similar to the best response correspondence, except that the function does not "jump" from one pure strategy to another. The difference is illustrated in Figure 8, where black represents the best response correspondence and the other colors each represent different smoothed best response functions. In standard best response correspondences, even the slightest benefit to one action will result in the individual playing that action with probability 1. In smoothed best response as the difference between two actions decreases the individual's play approaches 50:50.
There are many functions that represent smoothed best response functions. The functions illustrated here are several variations on the following function:
where represents the expected payoff of action , and is a parameter that determines the degree to which the function deviates from the true best response (a larger implies that the player is more likely to make 'mistakes').
There are several advantages to using smoothed best response, both theoretical and empirical. First, it is consistent with psychological experiments; when individuals are roughly indifferent between two actions they appear to choose more or less at random. Second, the play of individuals is uniquely determined in all cases, since it is a correspondence that is also a function. Finally, using smoothed best response with some learning rules (as in Fictitious play) can result in players learning to play mixed strategy Nash equilibria (Fudenberg & Levine 1998).
참고 항목
참조
- Ellison, G. (1993), "Learning, Local Interaction, and Coordination" (PDF), Econometrica, 61 (5): 1047–1071, doi:10.2307/2951493, JSTOR 2951493
- Fudenberg, D.; Levine, David K. (1998), The Theory of Learning in Games, Cambridge MA: MIT Press
- Fudenberg, Drew; Tirole, Jean (1991). Game theory. Cambridge, Massachusetts: MIT Press. ISBN 9780262061414. 책 미리보기.
- Gibbons, R. (1992), A primer in game theory, Harvester-Wheatsheaf, S2CID 10248389
- Nash, John F. (1950), "Equilibrium points in n-person games", Proceedings of the National Academy of Sciences of the United States of America, 36 (1): 48–49, Bibcode:1950PNAS...36...48N, doi:10.1073/pnas.36.1.48, PMC 1063129, PMID 16588946
- Osborne, M.J.; Rubinstein, Ariel (1994), A course in game theory, Cambridge MA: MIT Press
- Young, H.P. (2005), Strategic Learning and Its Limits, Oxford University Press
- Nisan, N.; Roughgarden, T.; Tardos, É.; Vazirani, V.V. (2007), Algorithmic Game Theory (PDF), New York: Cambridge University Press