back to top
Home NHSJS Reports Mathematical Models in Vestiaire: A Statistical Investigation into the Variables that Drive...

Mathematical Models in Vestiaire: A Statistical Investigation into the Variables that Drive Resale Outcomes

0
28

Abstract

This study investigates how one can predict how likely an item is to sell on Vestiaire Collective, a luxury clothing resale platform. Five statistical models were built and compared in R on a sample of 27,472 listings: linear regression, logistic regression, decision tree, random forest, and mixed effects. Four predictor variables were used: like count, product condition, brand name, and usual shipping time for the given seller. The models were evaluated using Root Mean Squared Error (RMSE). The mixed effects model performed best (RMSE = 0.446). Like count showed the strongest association with sale outcome: listings with zero likes had an observed sale probability of 11%, compared to 82% for listings with 50 or more likes. Product condition also had a statistically significant, though smaller, contribution to sale probability, with effects ranging from about 5 to 12 percentage points. These findings show that like count is strongly associated with sale outcomes on resale platforms, suggesting a meaningful role for social engagement signals in shaping buyer behavior.

Keywords: Mathematics; Statistics; resale prediction; social proof; Vestiaire Collective

Introduction

The market for secondhand and pre-owned luxury goods has grown substantially in recent years1, with platforms such as Depop, TheRealReal, and Vestiaire Collective connecting individual sellers with a global pool of buyers. This reflects a broader trend: the global secondhand apparel market exceeded $256 billion in 2025 and continues to expand rapidly2. This growth does not necessarily reflect a shift toward more sustainable consumption, however. Secondhand fashion consumption has been found to correlate positively with new clothing purchases rather than displacing them3, and resellers themselves cite financial gain and convenience, more than sustainability, as their primary motivations2. Consumer priorities also vary by market segment: luxury secondhand buyers value brand prestige and authenticity differently than the majority of buyers4, and often return to resale platforms for access to items unavailable through official brand channels5. Two reviews of secondhand consumption research each note a gap: most studies treat secondhand shopping as one general behavior rather than examining how specific platform features shape it6,7. Separately, other research has found that word-of-mouth and price are strong predictors of consumers’ intent to buy secondhand clothing8. Brand name may function similarly to seller reputation systems studied in other online marketplaces, which have been shown to reduce buyer uncertainty and support price premiums for trusted sellers9,10,11. Despite this abundance of research on consumer motivations, little quantitative research has evaluated which listing-level factors predict whether an individual item sells.

This project also examines social engagement signals that may relate to a sale. Prior work suggests visible engagement, such as like count, influences consumer decisions by signalling credibility and desirability12, an effect documented across social media shopping13 and short-form video commerce14. Online reviews function similarly, though inconsistent reviews have been shown to reduce purchase intention, particularly when purchasing for others15.

Beyond commercial applications, this project delves into the kinds of social engagement that eventually lead to a sale. Prior work suggests that visible engagement (such as like count in this case) influences consumer decisions by signalling credibility and desirability12.

This paper reports the construction and evaluation of five statistical models and the identification of like count as the dominant predictor of sale outcome on Vestiaire Collective.

Methods

Data Source and Collection

The dataset was obtained from Kaggle, a public data science platform16. It contains listing-level records from Vestiaire Collective, including product condition, brand name, like count, shipping time, price, and whether or not the item was sold. The data was publicly available and fully anonymised, containing no personally identifiable information about buyers or sellers, so no consent was required for its use. To ensure balanced class representation for model training, an equal number of sold and unsold listings (n = 13,736 each) were randomly sampled from the original data, spanning 2,512 distinct brands after merging duplicate brand names (e.g., “celine” and “Céline”). Data collection occurred at a single point in time; the exact scrape date was not recorded in the dataset’s metadata. An item was coded as sold (1) if the listing’s status field indicated it had been sold, and unsold (0) otherwise. Because this reflects a single snapshot rather than continuous tracking over time, we cannot fully distinguish confirmed non-sales from listings that had simply not yet sold at the time of collection; this limitation is discussed further in the Limitations section.

Data Cleaning and Feature Engineering

There were several cleaning steps to be done before using this data in statistical models. For instance, brand names that are known to be the same were merged into one (e.g. “celine” and “Céline”; “tiffany” and “Tiffany & Co.”). In total, 25 cleaned brand name categories were created by merging two or more raw name variants (e.g., 13 raw Nike variants and 12 Adidas variants), out of 2,471 distinct cleaned brands overall. Raw like counts were converted into categorical brackets (0; 1–5; 6–10; 11–15; 16–20; 21–30; 31–40; 41–50; 50+) to allow for non-linear effects through categorical encoding and to reduce noise from extreme values. To evaluate this choice, a logistic regression model using continuous like count was compared to the bracketed version by measuring how well each model fit the data. The bracketed model fit the data noticeably better, supporting the use of categorical brackets. Whether or not the item sold was turned into binary indicators (1 or 0). Of the 27,472 listings, 3,759 (13.68%) were missing shipping time data and were excluded, leaving 23,713 listings for modeling. Excluded listings had a modestly higher like count and price, but substantially lower sale rate(34.6% vs 52.4%), suggesting that missingness was not random with respect to outcome.

Four predictor variables were selected for final modelling based on theoretical relevance and performance across several tested variable combinations (Table 1). This comparison was conducted within the same cross-validation procedure used to report final model performance, rather than on the full dataset beforehand, to avoid inflating reported performance. Price was deliberately excluded from the four final predictor variables. Price was highly right-skewed, with a small number of listings, such as luxury watches, reaching values far higher than most of the dataset, risking disproportionate influence on coefficient estimates in the models. A supplementary check (see Discussion) was used to confirm that this exclusion does not leave an unaddressed confounding variable in the reported results.

VariableDescriptionRationale
Like BracketLikes grouped: 0, 1–5, 6–10, 11–15, 16–20, 21–30, 31–40, 41–50, 50+Direct measure of buyer interest
Product ConditionFair, Good, Very Good, Never Worn, Never Worn With TagSignal of quality
Brand NameCleaned to merge duplicates (e.g. celine / Céline)Many buyers are attracted to specific brand names
Ships WithinSeller dispatch window: 1–2, 3–5, 6–7, or 7+ daysShows the seller’s credibility and reliability
Table 1 | Predictor variables used in final models.

Statistical Models

Five models were constructed and evaluated. All were implemented in R17.

Linear regression models sale probability as a weighted sum of predictor variables. It is the most simple and interpretable approach but carries a fundamental limitation for binary outcomes: it can produce predictions outside the 0–1 range, which is impossible for probabilities. It was included as a baseline and for coefficient interpretability.

Logistic regression works similar to Linear regression, but addresses many of the limitations by converting the probabilities to a log-odds scale. It also better captures differences near the extremes of the probability range (e.g. the difference between 0.95 and 0.99 is not substantial in percentages for linear models but in log-odds they are 2.9 and 4.6).

A decision tree lets the model choose which variables it thinks are most important, instead of the user manually using intuition. It requires no assumptions about the functional form of predictor-outcome relationships. However, excessive splitting can lead to overfitting. The decision tree was fit with a maximum depth of 3 and no complexity-based trimming (cp = 0).

A random forest builds many decision trees and averages their predictions. This typically reduces overfitting relative to a single tree. In this study, however, the random forest had the worst performance. One possible explanation, discussed further below, is that its random subsampling of variables at each split limits how consistently it can draw on like count, the dominant predictor. The random forest was fit using 500 trees (default) with 4 variables considered at each split (mtry = 4).

A mixed effects model separates predictors into fixed effects — consistent predictors estimated across the full dataset (condition, like bracket) — and random effects, variables that can take on many different values, which can be difficult for the other models to account for (like brand name)18. Estimates for rare brands are given less weight, preventing overfitting. Formally, the model was specified as a binomial mixed-effects logistic regression with a logit link: sold ~ like_bracket + usually_ships_within + product_condition + (1 | brand_name_clean), fit using the glmer() function from the lme4 package. Brand (2,512 levels) was treated as a random intercept, letting brands with fewer listings shrink toward the overall average rather than overfit. The model converged without warnings, and the random intercept’s variance (0.233, SD = 0.482) showed meaningful differences in baseline sale likelihood across brands even after accounting for like count, condition, and shipping time — supporting the use of a random-effects structure for brand rather than treating it as a fixed effect.

Evaluation

All models were evaluated using Root Mean Squared Error (RMSE), a measure of how different the predicted outcome was from the actual outcome, computed via 20-fold cross-validation. This means that R split the dataset randomly into 20 groups (folds). For each of the 20 iterations, the model was trained on 19 of the folds and tested on the remaining held-out fold, with each fold serving as the test set exactly once. RMSE was then averaged across all 20 iterations. A lower RMSE value indicates better performance. RMSE was used as the primary metric because it works consistently across all five models, including linear regression, which doesn’t produce classification probabilities the same way. On a binary outcome, RMSE is also equivalent to the Brier score, a standard measure of how accurate predicted probabilities are. RMSE was used as the primary method for comparing models, since it applies consistently across linear and non-linear approaches. To more fully characterize classification performance, AUC, accuracy, sensitivity, specificity, precision, recall, and a calibration check were computed for the mixed effects and logistic regression models(see Results).

Results

Model Performance

Table 2 presents RMSE values for all five models. The mixed effects model achieved the lowest RMSE (0.446), so it was the best-performing model. Random forest performed worst (RMSE = 0.462).

ModelRMSENotes
Linear Regression0.449Baseline; interpretable but not suited to probability outcomes since it often predicts more than 100% chance of selling
Logistic Regression0.449Better suited to binary outcomes; Slightly outperformed linear regression (statistically significant, p = .006), despite similar RMSE
Decision Tree0.451Slightly higher than logistic
Random Forest0.462Worst performer; likely limited by inconsistent access to like count across trees (see Discussion)
Mixed Effects0.446Best performer; significantly outperformed logistic regression
Table 2 | Model RMSE values (20-fold cross-validation). Lower RMSE indicates better performance.

To check whether these differences in performance were meaningful, paired t-tests were used to compare cross-validated RMSE across all five models. Logistic regression performed significantly better than the decision tree and random forest (both p < .001), and also outperformed linear regression, though that difference was very small (p = .006). The mixed effects model outperformed logistic regression as well, with a significantly lower average RMSE (p < .001), confirming it as the strongest-performing model among those tested.

Classification Performance

The mixed effects model achieved AUC=0.757, accuracy=69.6%, sensitivity=72.5%, specificity=66.2%, precision=70.8%, and recall=72.5% on a held-out test split (logistic regression: AUC=0.749, accuracy=68.7%). A 10-bin calibration check showed close agreement between predicted and observed sale rates (eg. 83.2% predicted vs 83.3% observed), indicating reliable probabilities.

Like Count as the Dominant Predictor

Like count was by far the strongest predictor of sale outcome. Table 3 shows predicted sale probabilities by like bracket, derived from the linear regression model. Linear regression coefficients were used here because they translate directly into percentage-point changes in sale probability, which are more intuitive to communicate than the mixed effects model’s log-odds coefficients. The same overall pattern held in the mixed effects model: like-count coefficients increased sharply at low brackets and leveled off at higher ones, confirming that this finding is not an artifact of the choice of model.

Like BracketSale ProbabilityConfidence Interval
011%N/A
1–543%[30.6, 34.4]
6–1065%[51.8, 56.2]
11–1570%[57.2, 62.2]
16–2076%[62.6, 68.3]
21–3078%[64.0, 69.7]
31–4082%[67.9, 75.2]
41–5080%[64.9, 74.0]
50+82%[66.9, 75.3]
Table 3 | Sale probability assumes Fair condition, the baseline brand category, and 1–2 day shipping (the model’s reference levels). The 95% CI reflects uncertainty in the like-count effect itself, in percentage points relative to 0 likes.

The most striking result is the jump from zero to one-to-five likes: a gain of 32.7 percentage points. This suggests that early engagement has a disproportionately large effect on predicted sale probability. The relationship flattens substantially beyond 20 likes, with diminishing returns. The total range across all brackets is 71 percentage points, larger than the effect of any other variable in the model. One possible explanation, not directly tested here, is that buyers perceive items with more likes as more desirable or at greater risk of selling soon, creating an incentive to act quickly. This kind of herding behavior, in which visible signals of others’ interest shift individual purchase decisions, has been documented across a range of online marketplaces19,20,21,22,23, and related work has shown that scarcity and urgency cues can similarly accelerate purchase decisions24,25. Because the dataset reflects a single snapshot in time rather than continuous tracking, this finding should be interpreted as an association observed at time of scrape, rather than as evidence that early likes cause or predict a future sale.

Product Condition

Product condition surprisingly contributed modest effects relative to like count. Linear regression coefficients (relative to Fair condition as the baseline) are reported in Table 4.

Product ConditionCoefficient vs. Fair95% CIp-value
Fair0.00 (intercept)Baseline 
Good+0.049[.006, 0.091].024
Very Good+0.117[0.077, 0.157]<.001
Never Worn+0.090[0.049, 0.131]<.001
Never Worn With Tag+0.116[0.075, 0.158]<.001
Table 4 | All coefficients represent percentage-point change in sale probability relative to Fair condition, holding like count, brand, and shipping time constant.

All four non-baseline conditions showed a statistically significant positive effect relative to Fair condition (all p < .05). The effect size increased roughly in order of quality, with Good showing the smallest effect (+4.9 percentage points) and Very Good and Never Worn With Tag showing the largest, nearly identical effects (+11.7 and +11.6 percentage points). This aligns with broader research on online markets showing that buyers face persistent uncertainty about whether a product matches their expectations, particularly for goods that cannot be physically inspected before purchase26.

Discussion

The main finding of this study is the major role of like count in predicting sale outcomes. This is consistent with literature on social proof in consumer behaviour: prior engagement signals credibility and desirability to subsequent viewers12. As a check on whether the observed like-count effect could instead be attributable to price, a logistic regression model was fit with log-price added. Like-count coefficients increased across every bracket(eg., the 50+ likes coefficient rose from 2.98 to 4.31 in log odds), indicating price does not confound the central finding. Product category and seller country were available in the dataset but excluded to keep model concise.

The nonlinear shape of like-probability echoes the research on the role of zero-reviews versus any-review effects in online ratings systems. The steep initial gain (0 to 1–5 likes: +32 percentage points) and then diminishing returns point to the fact that a listing with any positive engagement is perceived extremely differently from one with none.

As proposed earlier, random forest’s weaker performance may stem from its random subsampling of variables at each split, which can limit the consistency of individual trees. To explore this, each predictor’s contribution to the random forest’s accuracy was examined. Like count stood out clearly as the most important variable, contributing far more than product condition, brand, or shipping time, cohering with the proposed explanation. As an additional check, the model’s error rate on unseen data was consistent with its cross-validated performance in Table 2, confirming the result is stable rather than specific to one data split.

Several limitations of this study should be noted. First, one should not assume a causation relationship between like count and sale. Items that would have sold for other reasons (brand, price, timing) may also accumulate likes, introducing confounding variables. Separating causation from correlation would require experimental data, such as a randomised promotion study. This study assessed the magnitude of potential confounding in two ways: the price-adjusted check described above (see Discussion) showed that the like-count effect did not weaken when price was added to the model, and the primary analysis of like count and product condition (Tables 3 and 4) controlled for brand and shipping time directly, rather than considering like count in isolation. Together, these checks suggest the like-count association is not simply a byproduct of the other listing characteristics available in this dataset, though other confounding factors, such as timing or seasonality, cannot be ruled out. The final dataset contained no missing values in variables(product condition, like count, brand name). Missingness in the original unsampled dataset was not separately assessed.

Multicollinearity among predictors was also assessed using variance inflation factors (VIF). All predictors showed low VIF values, indicating minimal correlation among like count, product condition, brand, shipping time, and price, and supporting the reliability of each variable’s individual coefficient estimates.

Second, the dataset reflects a snapshot in time rather than continuous tracking. As a result, listings coded as unsold may include both confirmed non-sales and items that were simply still active and had not yet sold at the time of collection. We cannot distinguish between these cases with the available data. Also, variables such as year and seasonal demand were not accounted for.

These findings are based on a single platform, Vestiaire Collective, and may not generalize directly to other resale marketplaces. Vestiaire specializes in authenticated luxury goods and has a like/engagement feature that may function differently than analogous features on other platforms (e.g., Depop, Poshmark, TheRealReal), which vary in audience, price range, and the visibility of engagement signals to buyers. The strength of the like-count association observed here should therefore be interpreted as specific to this platform and product category until tested elsewhere.

Conclusion

This study applied five statistical models to predict sale outcomes for listings on Vestiaire Collective. A mixed effects model had the best performance (RMSE = 0.446). Like count showed the strongest association with sale outcome, with observed sale rates rising from 11% at zero likes to 82% at 50 or more, and a 32 percentage point gain between zero and one-to-five likes. Product condition had a smaller but statistically significant effect on sale probability compared to like count.

These findings contribute to a broader understanding of how social engagement signals translate into economic outcomes on clothing resale sites.

Acknowledgments

I would like to thank my mentor, also my Advanced Statistics teacher, a graduate in Statistics, for guidance on mixed effects modelling and research design, and for feedback throughout the project.

References

  1. ThredUp. 2024 Annual Resale Report; ThredUp: San Francisco, CA, 2024. [↩]
  2. Herman, J.; Kim-Vick, J.; Hyun, J. The Inner Drive: Unpacking the Motivations for Consumer Participation as Sellers in Apparel Resale. Businesses Vol. 5, No. 4, pg. 53, 2025, https://doi.org/10.3390/businesses5040053. [↩] [↩]
  3. Mizrachi, M. P.; Sharon, O. Secondhand fashion consumers exhibit fast fashion behaviors despite sustainability narratives. Sci. Rep. Vol. 15, No. 1, 2025, https://doi.org/10.1038/s41598-025-19089-1. [↩]
  4. ul Hasan, H. M. R.; Lang, C.; Xia, S. Investigating Consumer Values of Secondhand Fashion Consumption in the Mass Market vs. Luxury Market: A Text-Mining Approach. Sustainability Vol. 15, No. 1, pg. 254, 2022, https://doi.org/10.3390/su15010254. [↩]
  5. Murtas, G.; Pedeliento, G. Investigating the Customer Journey in Second-Hand Fashion Platforms: Implications for Luxury Brand Management. J. Consum. Behav. Vol. 24, No. 2, 2024, https://doi.org/10.1002/cb.2442. [↩]
  6. Evans, F.; Grimmer, L.; Grimmer, M. Consumer orientations of secondhand fashion shoppers: The role of shopping frequency and store type. J. Retailing Consum. Serv. Vol. 67, No. 1, 102991, 2022, https://doi.org/10.1016/j.jretconser.2022.102991. [↩]
  7. Gilal, F. G.; Shaikh, A. R.; Yang, Z.; Gilal, R. G.; Gilal, N. G. Secondhand consumption: A systematic literature review and future research agenda. Int. J. Consum. Stud. Vol. 48, No. 3, 2024, https://doi.org/10.1111/ijcs.13059. [↩]
  8. Cuong, D. T. Examining How Factors Consumers’ Buying Intention of Secondhand Clothes via Theory of Planned Behavior and Stimulus Organism Response Model. J. Open Innov. Technol. Mark. Complex. Vol. 10, No. 4, 100393, 2024, https://doi.org/10.1016/j.joitmc.2024.100393. [↩]
  9. Jiao, R.; Przepiorka, W.; Buskens, V. Reputation effects in peer-to-peer online markets: A meta-analysis. Soc. Sci. Res. Vol. 95, 102522, 2021, https://doi.org/10.1016/j.ssresearch.2020.102522. [↩]
  10. Ba, S.; Pavlou, P. A. Evidence of the effect of trust building technology in electronic markets: Price premiums and buyer behavior. MIS Q. Vol. 26, No. 3, pg. 243–268, 2002, https://doi.org/10.2307/4132332. [↩]
  11. Standifird, S. S. Reputation and e-commerce: eBay auctions and the asymmetrical impact of positive and negative ratings. J. Manage. Vol. 27, No. 3, pg. 279–295, 2001, https://doi.org/10.1016/S0149-2063(01)00092-7. [↩]
  12. Cialdini, R. B. Influence: The psychology of persuasion; Harper Business: New York, NY, 2001. [↩] [↩] [↩]
  13. Talib, Y. Y.; Saat, R. Social proof in social media shopping: An experimental design research. SHS Web Conf. Vol. 34, 02005, 2017, https://doi.org/10.1051/shsconf/20173402005. [↩]
  14. Huang, W.; Wang, X.; Zhang, Q.; Han, J.; Zhang, R. Beyond likes and comments: How social proof influences consumer impulse buying on short-form video platforms. J. Retailing Consum. Serv. Vol. 84, 104199, 2024, https://doi.org/10.1016/j.jretconser.2024.104199. [↩]
  15. Wang, J.; Liu, Y.; Qiu, Z.; Zhao, Z.; Wang, M.; Lan, H. Shopping for others: how does inconsistency in online reviews affect purchase intentions differently? Front. Psychol. Vol. 16, 1579545, 2025, https://doi.org/10.3389/fpsyg.2025.1579545. [↩]
  16. Vestiaire Collective Listings Dataset. Kaggle. https://www.kaggle.com (accessed May 2026). [↩]
  17. R Core Team. R: A language and environment for statistical computing; R Foundation for Statistical Computing: Vienna, Austria, 2024. [↩]
  18. Bates, D.; Maechler, M.; Bolker, B.; Walker, S. lme4: Linear mixed-effects models using Eigen and S4. J. Stat. Softw. Vol. 67, pg. 1–48, 2015, https://doi.org/10.18637/jss.v067.i01. [↩]
  19. Luo, P.; Wu, B.; Law, R.; Xu, Y. Herding behavior in peer-to-peer trading economy: The moderating role of reviewer photo and name. Tour. Manage. Perspect. Vol. 45, 101050, 2023, https://doi.org/10.1016/j.tmp.2022.101050. [↩]
  20. Chen, Y.-F. Herd behavior in purchasing books online. Comput. Hum. Behav. Vol. 24, No. 5, pg. 1977–1992, 2008, https://doi.org/10.1016/j.chb.2007.08.004. [↩]
  21. Ali, M.; Amir, H. Understanding consumer herding behavior in online purchases and its implications for online retailers and marketers. Electron. Commer. Res. Appl. Vol. 64, 101356, 2024, https://doi.org/10.1016/j.elerap.2024.101356. [↩]
  22. Ali, M.; Amir, H.; Shamsi, A. Consumer Herding Behavior in Online Buying: A Literature Review. Int. Rev. Manage. Bus. Res. Vol. 10, No. 1, pg. 345–360, 2021, https://doi.org/10.30543/10-1(2021)-30. [↩]
  23. Sunder, S.; Kim, K. H.; Yorkston, E. A. What Drives Herding Behavior in Online Ratings? The Role of Rater Experience, Product Portfolio, and Diverging Opinions. J. Mark. Vol. 83, No. 6, pg. 93–112, 2019, https://doi.org/10.1177/0022242919875688. [↩]
  24. Ali, F.; Maqsood, D. H.; Janjua, D. Q. Psychological Triggers in Online Shopping: The Influence of Scarcity, Urgency, and Personalization on Consumer Buying Behavior. Crit. Rev. Soc. Sci. Stud. Vol. 3, No. 2, pg. 269–289, 2025, https://doi.org/10.59075/cxyapm95. [↩]
  25. Wrabel, A.; Kupfer, A.; Zimmermann, S. Being Informed or Getting the Product? How the Coexistence of Scarcity Cues and Online Consumer Reviews Affects Online Purchase Decisions. Bus. Inf. Syst. Eng. Vol. 64, No. 5, pg. 575–592, 2022, https://doi.org/10.1007/s12599-022-00772-w. [↩]
  26. Hong, Y.; Pavlou, P. A. Product fit uncertainty in online markets: Nature, effects, and antecedents. Inf. Syst. Res. Vol. 25, No. 2, pg. 328–344, 2014, https://doi.org/10.1287/isre.2014.0520. [↩]

LEAVE A REPLY

Please enter your comment!
Please enter your name here