Objective
The objective of this project was to evaluate the quality of the Job Posting Data in Hungary dataset obtained from Kaggle.com and assess its suitability for conducting a reliable analysis of the job market.
Project Status: The project documentation has been completed. The GitHub repository is currently being updated with SQL scripts and supporting source materials used during the data quality audit process.
Project: Job Posting Data Analysis
Tools: Power Query / Pandas: Profiling, Cleansing, and Data Preparation Sql: Completeness Analysis
ETL Process Design and Automation
Description: The project initially aimed to analyze the Hungarian job market using a job postings dataset. However, a data quality and representativeness assessment revealed that only 16 out of nearly 10,000 records were related to Hungary. Since such a limited sample could not support reliable conclusions, the project scope was revised to focus on the entire dataset.
Before the analysis, the data was cleaned and prepared in Python using Pandas. Low-value columns were removed, and the dataset structure was standardized. The refined dataset was then loaded into a database, where SQL was used to analyze recruitment requirements, in-demand skills, technologies, and other key characteristics of published job postings.
Dataset
The dataset was analyzed using the following variables:
Project Scope
As part of the project, a comprehensive data profiling and quality assessment was conducted on a job postings dataset. The work began with an exploration of the dataset structure and the identification of key attributes relevant to labor market analysis. SQL queries were then used to verify the consistency of location data, job titles, and occupational categories. Particular attention was given to missing values, inconsistent data formats, character encoding issues, and variations in the level of detail across records. The project also evaluated the analytical value of individual columns and identified limitations affecting the dataset's suitability for business analysis. The final stage involved drawing conclusions regarding the dataset's readiness for further analysis and reporting.
Key Findings and Conclusions
During the data profiling phase, several significant data quality issues were identified, substantially limiting the dataset's suitability for reliable job market analysis. The most critical problems were found in the location data, where countries, cities, administrative regions, incomplete entries, and incorrectly formatted values were stored within the same column. In addition, character encoding issues were detected, making it difficult to accurately identify certain locations.
The analysis of job titles revealed a high level of variation and a lack of standardized naming conventions. As a result, similar roles appeared under multiple title variations, making reliable aggregation and comparison impossible without prior data normalization. Similar issues were observed in the occupational categories, which were assigned inconsistently, while some records contained undefined or missing category values.
Among the attributes analyzed, only the Seniority variable exhibited a relatively consistent structure and a limited set of values. However, it is important to note that it reflected the level of managerial responsibility rather than a traditional measure of professional experience. As a result, its usefulness for labor market analysis was limited.
The data quality assessment led to the conclusion that the dataset was not sufficiently prepared for reliable job market analysis. Using the data without prior standardization and reconstruction of certain information could have resulted in inaccurate or misleading analytical conclusions.
Conclusion Conducting a reliable job market analysis based on the evaluated dataset was not feasible without extensive data normalization. The project concluded with the identification of key data quality issues and an assessment of the dataset’s readiness for further analysis.