tagboard

What you are actually building

Not a model — a dataset. The platform trains the same model, the same way, for everyone: 400 steps, one fixed seed, no early stopping. The only thing that differs between you and the team next to you is the rows you hand it. That is the whole exercise, and it is why a better score means you understood the data better, not that you got a luckier run.

A row is a claim about an article, and whether that claim is true:

{"article_id": "art-0001", "claim": "/categorie/sport/voetbal", "label": 1, "src": "positives_v1"}

The model that learns from them is a cross-encoder: it reads the claim and the article together, as one sequence, and answers one question — does this claim hold for this article? That is why it can judge a tag it has never seen paired with that article before, and why the wording of your negatives matters as much as their quantity.

The loop

most of your day
  1. Write build.py in the browser, on My log. It gets a corpus, a taxonomy and a seeded rng, and returns a list of rows.
  2. Run it. This is free and unlimited — it executes in a sandbox on the server, with no network, a 60-second wall clock and a memory cap. Run it as often as you like; it costs you nothing.
  3. Look at the rows it made. The inspector pages through them and lets you filter by src and label. This is the step people skip, and it is the step that finds the bug.
  4. Submit. This is the one that costs. You get 15 of them, and a submission is queued, trained and scored.
  5. Read the result — the loss curves, the per-category confusion, and the gap between your own validation and the public score.

Running is free; submitting is not. If you are unsure whether a change helped, the honest answer usually comes from looking at rows, not from spending a slot to find out.

Three numbers, and why they disagree

NumberMeasured onWhat it is for
Own validation 10% of your own rows, held out before training Your estimate. The platform measures it — you never report it yourself.
Public F1 A hidden set nobody can train on What the board ranks you by, most of the day.
Sealed F1 A second hidden set, revealed once, late The one that counts. Scored against your latest model, not your best.

The gap between the first two is the interesting one. An own-validation score far above your public F1 means your rows taught the model to do well on rows like yours — which is a warning, not an achievement.

The model you train

one knob, and it is not the interesting one

Every submission fine-tunes the same weights the same way — 400 steps, one fixed seed, no early stopping. The only thing you set is how much of each article the model is allowed to read.

OptionReadsRun Peak GPU
RobBERT 256 256 tokens · ~169 words 56s 3.2 GB Fastest. Sees roughly the first 170 words of an article.
RobBERT 512 512 tokens · ~338 words 96s 4.8 GB The default. Same weights, twice the window — the median article fits whole.

Cheaper is not worse until it costs you accuracy. The board breaks an exact tie in favour of the cheaper model, because two runs that score the same are not equally good — one of them used half the GPU and finished in half the time. That is the call you would have to defend in production.

Everything else about your run is the rows you hand it, which is the whole point of the exercise.

Measured on 400 steps · batch 16 · fp16 · one L4.

The number is noisier than it looks

Submit the same dataset twice and the score moves by about 0.02 on its own. Some of that is the seed; some is that the GPU's own arithmetic is not perfectly repeatable between runs.

So a change of ±0.01 is not a result. Before you conclude that an idea worked, ask whether it moved the score by more than the machine moves it by anyway.

What gets a dataset rejected

checked before a slot is spent

A rejected submission costs you nothing — you get the full list of problems and can fix them. These are the ones that are not about formatting:

Training on the hidden set
Articles held back for scoring are visible to your build.py but flagged in_hidden. Use one and the dataset is refused — not penalised, refused.
An ancestor used as a negative
If /categorie/sport/voetbal is true, then /categorie/sport is true as well. Labelling the parent 0 teaches the model something false.
A positive that is not true
A label: 1 on a tag the article does not carry, and which is not an ancestor of one it does.
Too many rows
3,000 at most. More rows is not the lever you think it is; the cap exists to make you choose.

corpus.is_true_for(claim, tags) answers the entailment question the same way the validator does — check before you emit, rather than finding out afterwards.

What build() is given

read-only · no network · no database
corpus.trainable Everything you may train on — the corpus minus the hidden set. Start here.
corpus[article_id] One article: .id, .text (Dutch prose), .true_tags.
for a in corpus: All of it, hidden articles included and flagged .in_hidden.
corpus.tags Every tag in the taxonomy — also passed to you as taxonomy.
corpus.ancestors(tag) Its parents, nearest first.
corpus.children(tag)
corpus.siblings(tag)
Down one level, and across.
corpus.is_true_for(claim, tags) Would the validator call this true?
rng A seeded random.Random. Same code, same dataset, every run.

Each row you return is {"article_id": …, "claim": …, "label": 0 or 1, "src": "…"}. src is free text naming the strategy that produced the row; the platform breaks your results down by it, so distinct names are how you find out which of your ideas actually paid.

The site you will test

this afternoon · /site

A newsroom archive of 60 articles, each showing the tags an editor supposedly gave it. Some of those tags are wrong. Nothing on the page says which — finding them is what your model is for.

Point Playwright at https://tagboard.jorenjanssens.be/site. Every element a test needs carries a data-testid, and those will not change under you:

SelectorWhereWhat it gives you
article-list/site the index, with data-article-count
article-link/site one per article, with data-article-id
article/site/articles/<id> the piece, with data-article-id
article-titlearticlethe headline
article-bodyarticle the prose, and nothing but the prose
tag-listarticle the chips, with data-tag-count
tagarticle one per tag. data-tag is the full taxonomy path, data-leaf and the text are the leaf — the leaf is what your model was trained on

https://tagboard.jorenjanssens.be/site/api/articles.json returns the whole archive in one request — id, title, url, text and tags — if you would rather enumerate than crawl. It carries exactly what the pages carry and no verdicts.

The site never says which tags are wrong, in the DOM or the JSON. There is an answer key; it is admin-only, and it is for the debrief.

The pages

My log
Where you work: the editor, your runs, the inspector, and every submission you have made with its curves and confusion.
Board
Everyone's best score, and the sealed result once it is published.
Playground
Paste an article, write any claim you like, and run your own trained model against it in your browser. The fastest way to find out what your model actually learned — and it takes no GPU time from anyone.