Twelve years with DiffVal and TDV: the making of a method
The post provided by Tiago Monteiro-Henriques

This post refers to the article TDV-optimization: A novel numerical method for phytosociological tabulation by Tiago Monteiro-Henriques published in Vegetation Classification and Survey (https://doi.org/10.3897/VCS.140466)
In 2007, during my PhD studies, my wife and I decided to raise our children in Campo Benfeito—a tiny village of about 50 inhabitants in a somewhat isolated Portuguese mountain region, and the land of my maternal lineage. Securing a position in Lisbon or in any other Portuguese university was seen as impossible at the time, and we wanted to stay in Portugal. Thus, located at the heart of my PhD study area, the village seemed quite fitting.
After I finished my PhD (2010), I maintained a tenuous connection with some universities, doing some research. My fascination with science kept calling me. I remember the long nights when I developed DiffVal and TDV, around 2013, in an old stone house at 1000 m altitude, the place of birth of my grandmother. I recall a huge mental struggle, trying to clarify the meaning of absences in vegetation science. I had the sensation that some established ideas—or simply semantics—could be blurring my reasoning: I intuited I was close to a simple, interesting calculation, but I was not able to see it clearly. In my trials and errors, I had a eureka moment when I found an index (DiffVal) that could get close to the old tabulation technique through optimization (TDV-optimization).
I presented DiffVal and TDV-optimization at the Ljubljana EVS meeting in 2014 (Monteiro-Henriques and Bellu 2014). The index calculations were shown, but the mathematical notation and the nomenclature I used—especially for the absences—was still undeveloped and probably not very convincing. At the time, it did not raise much interest. Some colleagues commented that so many numerical methods had been published that I would need to check whether something similar had already been proposed. The task of going through more than 100 years of documents seemed so intimidating that I left it dormant for some time.
For several years, I tested TDV-optimization using data from colleagues. The results were usually interesting but were rarely published, as the method itself remained unpublished.
From 2017 to 2023, I was funded by an FCT postdoctoral grant (Fundação para a Ciência e a Tecnologia), under the supervision of Professor Paulo Fernandes, Professor Jorge Orestes Cerdeira and Dr Mar Cabeza—I am very thankful for the full intellectual freedom they offered me. One of the grant’s objectives was to improve the computation of TDV-optimization and to make it properly available. A large part of the postdoc was dedicated to: (i) creating the diffval R package, (ii) developing new TDV-optimization procedures (linear integer programming and simulated annealing, with the indispensable support of Professor Orestes), and (iii) proving that optimizing TDV is an NP-hard problem—i.e., a complex computational problem with no known fast solution, where checking all possible answers could take an extremely long time (see also https://en.wikipedia.org/wiki/NP-hardness).
The package was released in 2022 and presented in a poster at the Madrid IAVS symposium that same year (Monteiro-Henriques et al. 2022). The new optimization procedures were included in the package and have recently been submitted for publication, together with Professor Orestes, along with the proof that the problem is NP-hard—a particularly demanding task accomplished by Professor Orestes.
At the end of the postdoc (2022-2023), within the fantastic working environment of the Global Change and Conservation research group (University of Helsinki), led by Dr Mar Cabeza, I delved into the review of historical literature on vegetation classification. Mar’s supervision provided me with the right stimuli and a peaceful environment that enabled me to further understand (and name) the two kinds of absences in the method: stochastic and differentiating absences.
I submitted a first version of the manuscript presenting DiffVal and TDV-optimization in July 2023. Professor David W. Roberts handled the manuscript, generously providing several constructive and relevant comments. Anticipating concerns that other researchers might raise, Professor Roberts sent me a series of questions to address. I had initially tried to avoid discussing these issues in the manuscript, as they risked reviving long-standing debates in vegetation science. However, I responded to the points he raised. Our respectful exchange of letters was extremely valuable in improving the article’s contextualization within vegetation science and in strengthening my arguments. Appendix 1 stemmed from this exchange. Moreover, the encouraging words that Professor Roberts consistently included in his letters gave me the morale to continue the endeavour.
In 2024, I was no longer receiving science funding. I dreamed of self-funding my research through agrarian activities with my wife (an oneiric vision some friends know I like to call Spartan-Arcadian science). I managed to acquire some original articles by Braun-Blanquet and by Ellenberg and wrote Appendix 1 (Terminological and epistemological considerations), but I admit that our family savings suffered. I thank my wife and three children for their understanding and love. I resubmitted an improved version of the manuscript in October 2024.

At the beginning of 2025, Professor Roberts asked me to address the comments of three anonymous reviewers: one with severe criticisms, another with more neutral feedback, and a third with positive remarks and a generous, monumental offering—a thorough grammatical correction of my systematic English errors, for which I am deeply thankful. The comments and criticisms led to a stronger manuscript, now including a meticulous comparison with other clustering methods and an improved discussion, culminating in the published version around twelve years after the creation of DiffVal and TDV. Professor Roberts’ acceptance letter was something very special to me and moved me deeply.
DiffVal and TDV are simple calculations with a specific aim: to identify clusters of objects that display patterns of exclusive attributes. Because these attributes are exclusive to one or a few clusters, they can be used to differentiate among them —that is, they function as differential attributes or, in vegetation science, as differential species.
The optimization process is slow, especially for large data sets, so, if possible, the user should consider dividing the data into smaller parts. Splitting the data set into physiognomic subsets is particularly advisable: since the indices are based on binary data, the method cannot identify vegetation units based on dominance patterns. Additionally, the user should keep in mind that differential species—as traditionally defined in phytosociological tabulation—are typically sought within a set of related relevés.
With this article, I hope to contribute to the advancement of vegetation science by conveying the following messages:
(i) General clustering methods overlook important requirements of traditional phytosociological tabulation. Moreover, reticulate patterns present additional challenges to most methods.
(ii) Statistical inference regarding species fidelity should be profoundly rethought. This approach can lead to the spurious identification of faithful species at high significance levels (α), and to the omission of truly faithful species at low α values.
(iii) The belief that presence-absence data are insufficient to replicate phytosociological tabulation, implying a loss of information, is disproven.
(iv) A perfect separation in ordination plots (e.g., nonmetric multidimensional scaling or correspondence analysis) does not guarantee that the underlying structure of the data set has been revealed.
(v) The mid-20th century shift from exclusive differential species to concentrated diagnostic species has profound consequences for vegetation classification, particularly as traditionally practised in phytosociology.
Looking back, I remain indebted to the people, places, and ideas that helped shape this work.
References:
- Monteiro-Henriques T, Bellu A (2014): An optimization approach to the production of differentiated tables based on new differentiability measures. In: Čarni A, Juvan N, Ribeiro D (Eds), Book of abstracts – 23rd International Workshop of the European Vegetation Survey. ZRC Publishing House, Ljubljana, Slovenia, 43–44. https://www.researchgate.net/publication/279956053_An_optimization_approach_to_the_production_of_differentiated_tables_based_on_new_differentiability_measures
- Monteiro-Henriques T, Cerdeira JO, Cabeza M, Fernandes PM (2022) A collection of R Tools for vegetation analyses. In: Abstracts book – 64th Annual Symposium of the International Association for Vegetation Science. IAVS, Madrid, Spain, 209. https://www.researchgate.net/publication/361936381_A_collection_of_R_Tools_for_vegetation_analyses
Brief personal summary of the author: Tiago Monteiro-Henriques is a researcher from Portugal, collaborating with the Global Change and Conservation research group at the University of Helsinki, Finland. His research focuses on data analysis, particularly in the field of vegetation classification. He is interested in modelling approaches that support native vegetation conservation.











