
Researchers from the AI and Digital Science Institute at the HSE Faculty of Computer Science have developed an approach for selecting recommendation algorithms more effectively. Their approach uses pairwise comparisons of algorithms to create a tournament table, with the overall ranking based on their performance across all datasets in the tournament. This can reduce the number of algorithms that need to be tested when developing new services, saving both time and money. The study was presented at the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD 2026).
Recommendation systems determine which products, films, songs, or publications to show users. To do this, they analyse users’ past behaviour and predict what they might be interested in next. For example, recommendation systems identify users with similar interests and take into account the sequence of their views or purchases.
However, there is no universal recommendation algorithm. A method that works well for an online store may not be suitable for an online cinema. Therefore, algorithms are usually pre-tested on existing datasets, and their performance is averaged across them. This approach, however, has its limitations: the resulting ranking does not account for the specifics of individual datasets and may be unstable, while testing all algorithms online with real users is costly and risky.
Researchers at HSE University have developed a methodology for comparing recommendation algorithms based on the Bradley–Terry model, which is used to rank competitors based on the outcomes of pairwise ‘matches.’ In this study, the recommendation algorithms were the ‘players,’ while the ‘matches’ were tests of the algorithms on different datasets. Two algorithms were compared using a selected metric, such as recommendation accuracy, and the one with the higher score was declared the winner. Based on the results of all pairwise comparisons, the model estimated the relative strength of each algorithm. The researchers trained and tested 14 algorithms on 89 datasets from different fields.
The results showed that the same algorithms could occupy different positions in the ranking depending on the type of data used in the ‘matches.’ For example, on sequential datasets, where the order of user actions is important, SASRec and GASATF ranked as the top performers. However, when a dataset did not contain an explicit sequence, these algorithms dropped to tenth and eleventh place, respectively, while LightGCN and ALS topped the ranking.
Additionally, the researchers tested the robustness of the rankings to incomplete data. The ranking produced by the Bradley–Terry model remained stable even when some comparisons were missing, meaning that not every algorithm was compared with every other algorithm. The authors also tested an extended version of the model on additional datasets, taking the context into account. In this case, the model had to predict the winners without directly comparing the recommendation algorithms.
'The idea of using a sports model came to us thanks to the work of our senior colleague Vladimir Spokoiny. To draw an analogy with a sports tournament, the outcome of a match is influenced not only by the players themselves but also by conditions such as the city or the weather. In our study, the characteristics of the dataset served as this context, including the number of users and items, the average length of users’ histories, and other parameters. If we train the model to take this context into account alongside the results of previous comparisons, it can use the characteristics of a new dataset to predict in advance which algorithms are likely to perform best,' says co-author Anton Lysenko, Expert at the International Laboratory of Stochastic Algorithms and High-Dimensional Inference.
In 78% of cases, the algorithm ranked first by the model was indeed among the top three. With the conventional approach of averaging performance metrics, this was the case in only 16% of cases. The authors attribute this difference to the fact that, unlike heuristic comparison methods, their model has a sound theoretical foundation, making the resulting rankings more reliable.
'Our approach helps identify which algorithms are best suited to a particular dataset before they are tested online. For example, if a bank needs to recommend loyalty programmes, it can feed the characteristics of a new dataset into the model to identify the most promising algorithms. Only these algorithms would then need to be tested on real users, rather than all available options. This saves both time and resources,' says co-author Sergey Samsonov, Head of the International Laboratory of Stochastic Algorithms and High-Dimensional Inference.
The study was carried out as part of a programme implemented by the HSE AI Research Centre and supported by a grant from the Ministry of Economic Development of the Russian Federation.

