Least-squares community extraction in feature-rich networks using similarity data

6533b7d3fe1ef96bd125ff79

RESEARCH PRODUCT

Least-squares community extraction in feature-rich networks using similarity data

Soroosh Shalileh Boris Mirkin Boris Mirkin

subject

Computer science Economics Kernel Functions Social Sciences 02 engineering and technology Least squares Infographics Translocation Genetic Geographical Locations Medical Conditions 0202 electrical engineering electronic engineering information engineering Medicine and Health Sciences Psychology Cluster Analysis Operator Theory Data Management Multidisciplinary Applied Mathematics Simulation and Modeling Q R Experimental Psychology Europe Feature (computer vision)Research Design Physical Sciences Medicine 020201 artificial intelligence & image processing Graphs Algorithms Network Analysis Network analysis Research Article Computer and Information Sciences Science Feature vector Scale (descriptive set theory)Research and Analysis Methods Column (database)Similarity (network science)020204 information systems Parasitic Diseases Least-Squares Analysis Feature data business.industry Data Visualization Biology and Life Sciences Pattern recognition Tropical Diseases Economic Analysis Malaria People and Places Artificial intelligence business Mathematics

description

We explore a doubly-greedy approach to the issue of community detection in feature-rich networks. According to this approach, both the network and feature data are straightforwardly recovered from the underlying unknown non-overlapping communities, supplied with a center in the feature space and intensity weight(s) over the network each. Our least-squares additive criterion allows us to search for communities one-by-one and to find each community by adding entities one by one. A focus of this paper is that the feature-space data part is converted into a similarity matrix format. The similarity/link values can be used in either of two modes: (a) as measured in the same scale so that one may can meaningfully compare and sum similarity values across the entire similarity matrix (summability mode), and (b) similarity values in one column should not be compared with the values in other columns (nonsummability mode). The two input matrices and two modes lead us to developing four different Iterative Community Extraction from Similarity data (ICESi) algorithms, which determine the number of communities automatically. Our experiments at real-world and synthetic datasets show that these algorithms are valid and competitive.

year	journal	country	edition	language
2021-07-01	PLoS ONE

10.1371/journal.pone.0254377 http://europepmc.org/articles/PMC8282089