PhD Studentship: Statistical Complexity in the Evolution of Biological Networks Advancing network science to understand the organisation, function and evolution of biological systems We are seeking an ambitious PhD student to develop a new generation of approaches for understanding the structure and evolution of biological networks. Biological systems are extraordinarily complex. Protein-protein interaction networks, gene regulatory networks, metabolic networks and other molecular interaction systems contain enormous numbers of components and interactions, and their organisation is both far from random and far from regular. Their networks exhibit hierarchy (a few nodes with many connections and many nodes with few), homophily (nodes which are more “similar” are more likely to connect), modularity and other forms of structure that emerge from the underlying biological processes governing how components interact. But only recently have methods been proposed to measure network complexity directly and parsimoniously. So, how statistically complex are biological networks? What mechanisms generate this complexity? And, critically, how is this network complexity related to biological function and evolutionary history? This PhD will address these questions by bringing together computational and evolutionary biology with network science and statistical physics. The project is aligned with SPNDL a Scottish Enterprise High Growth Spin-Out project developing an app for cross-species protein-protein interaction network visualisation. It is expected there will be significant interaction with this project, allowing the student access to a cutting-edge technology for studying and analysing PPI networks while providing new analytics that can add to the platforms utility. Research themes The project will be organised around three closely connected themes. 1. The statistical complexity of protein-protein interaction networks and other biological systems The first goal will be to establish how statistical complexity manifests itself in biological networks across species and taxonomical categories, with an initial focus on protein-protein interaction (PPI) networks. You will extend and apply normalised measures of statistical complexity to large-scale PPI networks (e.g. via STRING and related databases) across species, tissues, and conditions. 2. The roles of "popularity" and "similarity" in explaining statistical complexity Popularity (degree heterogeneity/node fitness) and similarity (latent geometric/functional proximity) have each been proposed as drivers of real-world network structure, and their combination has been shown to generate complexity levels matching real networks better than either alone. This theme will test and refine popularity–similarity generative models specifically for biological data, asking which combinations of hierarchy and geometry — and which link-formation mechanisms (e.g. common-neighbour/triadic-closure-like growth) — best explain the complex structure of biological networks, and what this implies about the hidden "similarity spaces" (functional, co-expression-based, structural) organising protein interactions. This could ultimately lead to models that explain why biological networks have the structures that they do, rather than simply describing those structures after the fact. 3. Evolutionary evidence linking statistical complexity and functional heterogeneity Moving from structure to consequence: does higher statistical complexity confer greater resilience, evolvability, or functional diversification? Building on evidence that gene-expression-based (rather than random or purely degree-based) attachment mechanisms enhance the resilience of PPI networks as new proteins are added, this theme will investigate whether statistically complex network architectures are evolutionarily favoured- comparing network growth models against real evolutionary data across species, and asking whether complexity itself is a signature (or driver) of functional heterogeneity and adaptive capacity. Possible questions here include: - Does network complexity increase, decrease or reorganise during evolution? - Are more statistically complex networks associated with greater functional diversity? - Are highly heterogeneous networks more evolvable, more robust, or both? - Do conserved and rapidly evolving proteins occupy systematically different positions in the complexity landscape? A particularly exciting possibility is to investigate whether statistical complexity provides a quantitative bridge between network structure and biological function: rather than viewing complexity as merely a mathematical property of a network, can we show that it reflects the functional heterogeneity and evolutionary history of the biological system? There is substantial scope for the project to evolve as interesting results emerge. We are particularly interested in students who want to develop new theory as well as apply existing methods, rather than treating biological networks simply as datasets to analyse. The ideal candidate We are looking for someone who is ambitious, intellectually curious and technically strong, with the independence to pursue difficult questions and the creativity to develop new ones. You will be excited by open-ended, foundational questions about biology and complex systems and willing to work across disciplinary boundaries. You might come from mathematics, physics, computer science, computational biology, bioinformatics or a related discipline. A strong quantitative background is highly advantageous; as is experience with biology and biological data. No single candidate will arrive expert in all of network science, complexity theory, and evolutionary biology — what matters most is aptitude, motivation, and a willingness to learn fast across these areas. Background Literature The project builds particularly on the following work: - Smith & Smith, Statistical complexity of heterogeneous geometric networks, PLOS Complex Systems (2025), which develops a normalised measure of statistical complexity and investigates the interaction between network heterogeneity, geometry and link-growth mechanisms. - Smith, Explaining the emergence of complex networks through log-normal fitness in a Euclidean node similarity space, Scientific Reports (2021), which develops the surface–depth framework combining node fitness/popularity with latent similarity and demonstrates its explanatory power across more than 100 real-world networks. - Klein et al., A computational exploration of resilience and evolvability of protein–protein interaction networks, Communications Biology (2021), which provides a foundation for investigating how mechanisms of network growth relate to resilience and evolution in PPI networks. These papers should be regarded as a starting point for further enquiry. The successful candidate will be expected to develop and refine the research questions during the PhD. What you'll gain - Training in cutting-edge network science methodology, from measure development and proof techniques to large-scale computational modelling - Direct engagement with major biological network resources (e.g. STRING, WormNet-derived interactomes) and cross-species comparative datasets - The opportunity to develop genuinely new theory addressing long-standing open questions about complexity and evolution - A collaborative, interdisciplinary supervisory environment spanning complex systems, data science, and evolutionary/computational biology - Opportunities to present at international conferences and publish in leading complex systems and biology journals Essential Criteria · A strong undergraduate or Master’s degree (1st or 2:1 equivalent) in Computer Science, Mathematics, Evolutionary Biology, Computational Biology, Physics or a related quantitative field. Desirable Criteria - Prior experience with biological databases (e.g., STRING, BioGRID, GEO). - Strong coding skills in Python, Julia, R, or C++, with experience handling network libraries (e.g., NetworkX, igraph) and large-scale biological datasets. - Experience in publishing peer-reviewed research or presenting at academic conferences. Funding & Environment Duration: 3 years (Full-time). Funding: Covers tuition fees at Home rates plus a standard, tax-free stipend aligned with research council rates. Research Environment: You will join the NeuraSearch Laboratory at Strathclyde, an interdisciplinary, supportive research team with strong international collaborations across data science applied to biomedical data. How to Apply Interested candidates should send an initial informal inquiry to Dr Keith Malcolm Smith at keith.m.smith@strath.ac.uk containing: 1. Curriculum Vitae (CV) including academic transcripts and contact details for two references. 2. Expression of interest (max 1 page) highlighting your suitability for the role, technical background, and a brief statement of interest in how you envision tackling the problems outlined. (to subscribe/unsubscribe the EvolDir send mail to evoldir@evoldir.net)