Die KI-Verordnung trat am 1. August 2024 in Kraft und gilt gestaffelt. Die Verbote und die KI-Kompetenzpflicht gelten seit Februar 2025, die Pflichten für KI-Modelle mit allgemeinem Verwendungszweck seit August 2025 und das Kernregime für Hochrisiko-KI seit August 2026; in Produkte eingebettete Hochrisikosysteme folgen 2027.
Demos run a short happy path once with a human watching. Production runs thousands of variations unattended, where per-step error rates compound and an unbounded permission scope turns a wrong decision into an incident. The deployments that work narrow the scope, verify each step cheaply, and gate every irreversible action.
Price per token has fallen sharply through better hardware, smaller distilled models and serving optimisations. Consumption has grown faster: longer contexts, reasoning models that generate far more tokens per answer, and agents that turn one user action into dozens of model calls. Falling unit prices with rising unit counts produce larger bills.
Answer engines synthesise a response from several sources and show it above or instead of the traditional link list, so a query that once produced a visit can now be resolved without one. Publishers see impressions and citations rise while click-through falls, which breaks the advertising model that assumed every answer required a page view.
On common benchmarks the best open-weight models now sit close to frontier commercial models, and for many routine tasks the difference is not noticeable. The remaining gaps show up in long-horizon reliability, tool use, very long contexts and safety tuning — and in the operational work of running them yourself.