Demszky Gábor, who served as mayor of Budapest from 1990 to 2010, wrote on his social media that he asked a generative artificial intelligence to compare the mayoral performances of “the three of us” — himself, Tarlós István and Karácsony Gergely. Tarlós István served as mayor from 2010 to 2019; Karácsony Gergely has been mayor since 2019.
The AI ranking and Demszky’s response
According to the results Demszky published, the AI returned scores on a 100-point scale with Tarlós István first at 78 points. Demszky received 77 points and Karácsony 75. Demszky noted that the differences were not large, but each mayor had strengths in different areas.
He also argued that such a straightforward comparison can be misleading: Tarlós, he said, benefited from eight years of government support, while Karácsony has governed against headwinds — a condition Demszky said largely applied to himself as well. For that reason, Demszky overrode the AI’s ranking in his own assessment, declaring that he should be considered first because, in his view, he had the greatest historical impact on the development of Budapest. He summarized the three tenures as follows: he changed Budapest the most, Tarlós was strong at carrying out large investments and financial consolidation, and Karácsony excelled in promoting a liveable, green, pedestrian- and bicycle-friendly Budapest.
Newsroom test and AI variability
The Telex newsroom ran the same question in Google AI mode and received a different result on a ten-point scale: Tarlós 7.5, Demszky 5.5, Karácsony 5. A Telex Techtud expert commented that such variation is to be expected: generative AI does not always produce the same output for the same prompt, nor can it be expected to generate perfectly repeatable answers over time.
Why context matters
The episode underscores that evaluating political performance is complex. Numerical scores depend on starting assumptions, weighting, and which sources, periods or contextual factors are emphasized — whether the evaluator is human or algorithmic. Demszky’s intervention illustrates that even when an AI produces a ranked list, political actors and the public may weigh historical impact and institutional conditions differently.



