What it does
The test uses fifty examples per task, running once in a single region. The cascade threshold was chosen using the same test data, and users must download the original dataset text themselves. A later update adds comparisons with a local model called Laya.