Search results for "Outlier"
showing 10 items of 73 documents
Using Complex Surveys to Estimate theL1-Median of a Functional Variable: Application to Electricity Load Curves
2012
Mean proles are widely used as indicators of the electricity consumption habits of customers. Currently, Electricit e De France (EDF), estimates class load proles by using point-wise mean function. Unfortunately, it is well known that the mean is highly sensitive to the presence of outliers, such as one or more consumers with unusually high-levels of consumption. In this paper, we propose an alternative to the mean prole: the L1-median prole which is more robust. When dealing with large datasets of functional data (load curves for example), survey sampling approaches are useful for estimating the median prole and avoid storing all of the data. We propose here estimators of the median trajec…
A new method to "clean up" ultra high-frequency data
2007
In the applied econometrics, the availability of ultra high-frequency databases is having an important impact on the research market microstructure theory. The ultra high-frequency databases contain detailed reports of all the financial market activity information which is available. However, ultra high-frequency databases cannot be directly used. On one hand recording mistakes can be present, on the other hand missing information has to be inferred from the available data. In this paper, we propose a simple method in order to clean up the ultra high-frequency data from possible errors and we examine the method efficacy when we analyze data by using an autoregressive conditional duration (A…
Efficient Dense Disparity Map Reconstruction using Sparse Measurements
2018
International audience; In this paper, we propose a new stereo matching algorithm able to reconstruct efficiently a dense disparity maps from few sparse disparity measurements. The algorithm is initialized by sampling the reference image using the Simple Linear Iterative Clustering (SLIC) superpixel method. Then, a sparse disparity map is generated only for the obtained boundary pixels. The reconstruction of the entire disparity map is obtained through the scanline propagation method. Outliers were effectively removed using an adaptive vertical median filter. Experimental results were conducted on the standard and the new Middlebury datasets show that the proposed method produces high-quali…
An approach to identify new antihypertensive agents using Thermolysin as model: In silico study based on QSARINS and docking
2019
Thermolysin is a bacterial proteolytic enzyme, considered by many authors as a pharmacological and biological model of other mammalian enzymes, with similar structural characteristics, such as angiotensin converting enzyme and neutral endopeptidase. Inhibitors of these enzymes are considered therapeutic targets for common diseases, such as hypertension and heart failure. In this report, a mathematical model of Multiple Linear Regression, for ordinary least squares, and genetic algorithm, for selection of variables, are developed and implemented in QSARINS software, with appropriate parameters for its fitting. The model is extensively validated according to OECD standards, so that its robust…
The Performance of the Gradient-Like Influence Measure in Generalized Linear Mixed Models
2015
A gradient-like statistic, recently introduced as an influence measure, has been proven to work well in large sample, thanks to its asymptotic properties. In this work, through small-scale simulation schemes, the performance of such a diagnostic measure is further investigated in terms of concordance with the main influence measures used for outlier identification. The simulation studies are performed by using generalized linear mixed models (GLMMs).
Performances of neural networks for deriving LAI estimates from existing CYCLOPES and MODIS products
2008
International audience; This paper evaluates the performances of a neural network approach to estimate LAI from CYCLOPES and MODIS nadir normalized reflectance and LAI products. A data base was generated from these products over the BELMANIP sites during the 2001-2003 period. Data were aggregated at 3 km x 3 km, resampled at 1/16 days temporal frequency and filtered to reject outliers. VEGETATION and MODIS reflectances show very consistent values in the red, near infrared and short wave infrared bands. Neural networks were trained over part of this data base for each of the 6 MODIS biome classes to retrieve both MODIS and CYCLOPES LAI products. Results show very good performances of neural …
Anomaly Detection in Dynamic Social Systems Using Weak Estimators
2009
Anomaly detection involves identifying observationsthat deviate from the normal behavior of a system. One ofthe ways to achieve this is by identifying the phenomena thatcharacterize “normal” observations. Subsequently, based on thecharacteristics of data learned from the “normal” observations,new observations are classified as being either “normal” or not.Most state-of-the-art approaches, especially those which belongto the family parameterized statistical schemes, work under theassumption that the underlying distributions of the observationsare stationary. That is, they assume that the distributions thatare learned during the training (or learning) phase, thoughunknown, are not time-varyin…
A comparison of different methods for speleothem age modelling
2012
Abstract Speleothems, such as stalagmites and flowstones, can be dated with unprecedented precision in the range of the last 650,000 a by the 230Th/U-method, which is considered as one of their major advantages as climate archives. However, a standard approach for the construction of speleothem age models and the estimation of the corresponding uncertainty has not been established yet. Here we apply five age modelling approaches ( StalAge , OxCal, a finite positive growth rate model and two spline-based models) to a synthetic speleothem growth model and two natural samples. All data sets contain problematic features such as outliers, age inversions, large and abrupt changes in growth rate a…
CLUSTERING INCOMPLETE SPECTRAL DATA WITH ROBUST METHODS
2018
Abstract. Missing value imputation is a common approach for preprocessing incomplete data sets. In case of data clustering, imputation methods may cause unexpected bias because they may change the underlying structure of the data. In order to avoid prior imputation of missing values the computational operations must be projected on the available data values. In this paper, we apply a robust nan-K-spatmed algorithm to the clustering problem on hyperspectral image data. Robust statistics, such as multivariate medians, are more insensitive to outliers than classical statistics relying on the Gaussian assumptions. They are, however, computationally more intractable due to the lack of closed-for…
A Combined Multi-Cohort Approach Reveals Novel and Known Genome-Wide Selection Signatures for Wool Traits in Merino and Merino-Derived Sheep Breeds.
2019
Merino sheep represents a valuable genetic resource worldwide. In this study, we investigated selection signatures in Merino (and Merino-derived) sheep breeds using genome-wide SNP data and two different approaches: a classical F-ST-outlier method and an approach based on the analysis of local ancestry in admixed populations. In order to capture the most reliable signals, we adopted a combined, multi-cohort approach. In particular, scenarios involving four Merino breeds (Spanish Merino, Australian Merino, Chinese Merino, and Sopravissana) were tested via the local ancestry approach, while nine pair-wise breed comparisons contrasting the above breeds, as well as the Gentile di Puglia breed, …