Every model in every election was put through the same pipeline: it voted
on all the Wahl-O-Mat statements before seeing any party answers or justifications. Then it
graded the parties' written justifications. This pipeline is identical across the
elections.
Step 1 · the corpus
What the parties actually wrote
Statement here always means one of the Wahl-O-Mat's Thesen: a single
proposition put to every party on the ballot. For each statement, every party gives a
position — agree, neutral or disagree — and a short written justification for it. All
justifications on this page are taken from the tool's own
data. These German texts are what the LLMs later grade
().
The English on this page — a translation made by a separate model — is presentation
only: the graders always read the German original.
Parties that left more than a quarter of the statements without a justification are
excluded from further analysis of graded scores.
Step 2 · voting
The models' voting positions
The model gets the entire Wahl-O-Mat questionnaire at once (only the statement
wording, no party names, no positions, no justifications) and responds with an agree,
neutral or disagree answer for every statement.
This is done once per model.
The model is also asked to pick up to a third of the statements to count double, and
to leave at most a third of its answers neutral. A sheet that breaks the neutral rule is
rejected and re-asked up to six times, and dropped if the model still will not commit.
That answer sheet is then played through the real Wahl-O-Mat, with every party
selected, so the voting match percentages on this page come from the
Wahl-O-Mat's own method. Except for when explicitly noted, the unweighted
Wahl-O-Mat percentages are used for any analysis and stats.
Step 3 · blinding and grading
The models' grading scores
Before any grading is done, party identity is stripped out of the
justifications.1
The models see the justifications of all parties for a single Wahl-O-Mat statement at
a time, in shuffled order. The context of one statement does not cross into the context
of another.
The model gives a grade and a reason for that grade for every justification (the
reason first, so that the grade is not rationalised after the fact). The motivation for
the grading is given to the model as one question:
Does reading this make me more or less willing to back this party on this
statement?
+2 much more willing, +1 somewhat, 0 no difference,
−1 somewhat less, −2 much less. Beyond that there is no checklist of what
should count for or against, and no worked examples.