УДК 636:004.8
E.G. Skvortsova, E.A. Skvortsov
Abstract. The article examines the problem of synthetic data generation as a key direction in the development of big data technologies in livestock farming. The main causes of high-quality labeled data scarcity are analyzed: high collection costs, ethical limitations, rarity of pathological conditions, and non-compliance with FAIR principles. Modern methods of synthetic data generation are systematized, including generative adversarial networks (GANs), variational autoencoders (VAEs), probabilistic methods based on copula functions, and rank-based approaches. Special attention is paid to architectural solutions for preserving correlation structures and ensuring biological plausibility of generated data. Based on the analysis of scientific literature, the main challenges are identified: the problem of model specificity to distribution types, risks of introducing artifactual relationships, and the need for multilevel quality control. Promising research directions are determined, including hybrid approaches combining the advantages of procedural generation and deep learning, as well as the development of specialized metrics for assessing synthetic data quality in livestock farming.
Keywords: synthetic data, big data, livestock farming, generative adversarial networks, variational autoencoders, data scarcity, GAN, VAE, quality control, correlation structures.
Download the full text of the article






