What you are actually building
Not a model — a dataset. The platform trains the same model, the same way, for everyone: 400 steps, one fixed seed, no early stopping. The only thing that differs between you and the team next to you is the rows you hand it. That is the whole exercise, and it is why a better score means you understood the data better, not that you got a luckier run.
A row is a claim about an article, and whether that claim is true:
{"article_id": "art-0001", "claim": "/categorie/sport/voetbal", "label": 1, "src": "positives_v1"}
The model that learns from them is a cross-encoder: it reads the claim and the article together, as one sequence, and answers one question — does this claim hold for this article? That is why it can judge a tag it has never seen paired with that article before, and why the wording of your negatives matters as much as their quantity.
The loop
most of your day-
Write
build.pyin the browser, on My log. It gets acorpus, ataxonomyand a seededrng, and returns a list of rows. - Run it. This is free and unlimited — it executes in a sandbox on the server, with no network, a 60-second wall clock and a memory cap. Run it as often as you like; it costs you nothing.
-
Look at the rows it made. The inspector pages through them and lets
you filter by
srcand label. This is the step people skip, and it is the step that finds the bug. - Submit. This is the one that costs. You get 15 of them, and a submission is queued, trained and scored.
- Read the result — the loss curves, the per-category confusion, and the gap between your own validation and the public score.
Running is free; submitting is not. If you are unsure whether a change helped, the honest answer usually comes from looking at rows, not from spending a slot to find out.
Three numbers, and why they disagree
| Number | Measured on | What it is for |
|---|---|---|
| Own validation | 10% of your own rows, held out before training | Your estimate. The platform measures it — you never report it yourself. |
| Public F1 | A hidden set nobody can train on | What the board ranks you by, most of the day. |
| Sealed F1 | A second hidden set, revealed once, late | The one that counts. Scored against your latest model, not your best. |
The gap between the first two is the interesting one. An own-validation score far above your public F1 means your rows taught the model to do well on rows like yours — which is a warning, not an achievement.
The model you train
one knob, and it is not the interesting oneEvery submission fine-tunes the same weights the same way — 400 steps, one fixed seed, no early stopping. The only thing you set is how much of each article the model is allowed to read.
| Option | Reads | Run | Peak GPU | |
|---|---|---|---|---|
| RobBERT 256 | 256 tokens · ~169 words | 56s | 3.2 GB | Fastest. Sees roughly the first 170 words of an article. |
| RobBERT 512 | 512 tokens · ~338 words | 96s | 4.8 GB | The default. Same weights, twice the window — the median article fits whole. |
Cheaper is not worse until it costs you accuracy. The board breaks an exact tie in favour of the cheaper model, because two runs that score the same are not equally good — one of them used half the GPU and finished in half the time. That is the call you would have to defend in production.
Everything else about your run is the rows you hand it, which is the whole point of the exercise.
Measured on 400 steps · batch 16 · fp16 · one L4.
The number is noisier than it looks
Submit the same dataset twice and the score moves by about 0.02 on its own. Some of that is the seed; some is that the GPU's own arithmetic is not perfectly repeatable between runs.
So a change of ±0.01 is not a result. Before you conclude that an idea worked, ask whether it moved the score by more than the machine moves it by anyway.
What gets a dataset rejected
checked before a slot is spentA rejected submission costs you nothing — you get the full list of problems and can fix them. These are the ones that are not about formatting:
- Training on the hidden set
- Articles held back for scoring are visible to your
build.pybut flaggedin_hidden. Use one and the dataset is refused — not penalised, refused. - An ancestor used as a negative
- If
/categorie/sport/voetbalis true, then/categorie/sportis true as well. Labelling the parent 0 teaches the model something false. - A positive that is not true
- A
label: 1on a tag the article does not carry, and which is not an ancestor of one it does. - Too many rows
- 3,000 at most. More rows is not the lever you think it is; the cap exists to make you choose.
corpus.is_true_for(claim, tags) answers the entailment question
the same way the validator does — check before you emit, rather than finding
out afterwards.
What build() is given
read-only · no network · no database| corpus.trainable | Everything you may train on — the corpus minus the hidden set. Start here. |
| corpus[article_id] | One article: .id, .text (Dutch prose),
.true_tags. |
| for a in corpus: | All of it, hidden articles included and flagged
.in_hidden. |
| corpus.tags | Every tag in the taxonomy — also passed to you as
taxonomy. |
| corpus.ancestors(tag) | Its parents, nearest first. |
| corpus.children(tag) corpus.siblings(tag) |
Down one level, and across. |
| corpus.is_true_for(claim, tags) | Would the validator call this true? |
| rng | A seeded random.Random. Same code, same dataset, every
run. |
Each row you return is
{"article_id": …, "claim": …, "label": 0 or 1, "src": "…"}.
src is free text naming the strategy that produced the row; the
platform breaks your results down by it, so distinct names are how you find out
which of your ideas actually paid.
The site you will test
this afternoon · /siteA newsroom archive of 60 articles, each showing the tags an editor supposedly gave it. Some of those tags are wrong. Nothing on the page says which — finding them is what your model is for.
Point Playwright at https://tagboard.jorenjanssens.be/site. Every element a test
needs carries a data-testid, and those will not change under you:
| Selector | Where | What it gives you |
|---|---|---|
| article-list | /site | the index, with data-article-count |
| article-link | /site | one per article, with data-article-id |
| article | /site/articles/<id> | the piece, with data-article-id |
| article-title | article | the headline |
| article-body | article | the prose, and nothing but the prose |
| tag-list | article | the chips, with data-tag-count |
| tag | article | one per tag. data-tag is the full taxonomy path,
data-leaf and the text are the leaf —
the leaf is what your model was trained on |
https://tagboard.jorenjanssens.be/site/api/articles.json returns the whole archive
in one request — id, title, url, text and tags — if you would rather enumerate
than crawl. It carries exactly what the pages carry and no verdicts.
The site never says which tags are wrong, in the DOM or the JSON. There is an answer key; it is admin-only, and it is for the debrief.
The pages
- My log
- Where you work: the editor, your runs, the inspector, and every submission you have made with its curves and confusion.
- Board
- Everyone's best score, and the sealed result once it is published.
- Playground
- Paste an article, write any claim you like, and run your own trained model against it in your browser. The fastest way to find out what your model actually learned — and it takes no GPU time from anyone.