This commit is contained in:
Joaquin Gottlebe
2025-07-30 15:57:48 +02:00
parent e49fcfac46
commit a639c34cee
273 changed files with 20151 additions and 0 deletions
@@ -0,0 +1,3 @@
__pycache__
.ipynb_checkpoints/
venv/
@@ -0,0 +1,14 @@
Copyright (C) 2025 Joaquin Gottlebe
This program is free software: you can redistribute it and/or modify
it under the terms of the GNU General Public License as published by
the Free Software Foundation, either version 3 of the License, or
(at your option) any later version.
This program is distributed in the hope that it will be useful,
but WITHOUT ANY WARRANTY; without even the implied warranty of
MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
GNU General Public License for more details.
You should have received a copy of the GNU General Public License
along with this program. If not, see <https://www.gnu.org/licenses/>.
@@ -0,0 +1,119 @@
# Exploratory Data Analysis on Employment in Germany (1999–2024)
**Version:** v1.0
**Description:**
This notebook presents an exploratory data analysis of employment trends in Germany from 1999 to 2024. Using official data from the German Federal Statistical Office (Destatis), we investigate key patterns in employment over time, with a focus on overall employment trends, the distribution between employment types (employees vs. self-employed), sectoral developments, and structural shifts between economic sectors. The objective is to uncover long-term transformations in the German labor market and highlight significant historical developments.
## 1. Overview
- **Purpose:** This project includes a Jupyter Notebook and a Python script for exploratory data analysis on employment in Germany.
- **Research context:** Developed as coursework for the "Research Software Engineering" module.
- **Notable features:** Analysis includes sectoral comparisons, employment type trends, and structural economic transitions over 25 years.
- **Date of creation:** 14.04.2025
## 2. Project Structure
- **Main folders:**
- `data/` – Contains raw and processed datasets (.csv)
- `docs/` – Project markdown description
- `src/` – Python scripts and notebooks (.py, .ipynb)
- **File formats used:** `.py`, `.ipynb`, `.md`, `.csv`, `.txt`
- **Project size:** ~230 MB
## 3. Installation
### Prerequisites
- Python 3.13.3
- Packages listed in `requirements.txt`
### Clone the repository
With SSH (Needs SSH key setup):
With HTTPS:
```bash
git clone https://gitup.uni-potsdam.de/gottlebe/eda-employment-germany.git
```
With SSH (Needs SSH key setup):
```bash
git clone git@gitup.uni-potsdam.de:gottlebe/eda-employment-germany.git
```
### Change directory
```bash
cd eda-employment-germany
```
### Environment setup (Optional)
The virtual environment is optional but recommended to avoid package conflicts.
For Linux/MacOS:
```bash
python3 -m venv venv
source venv/bin/activate
```
For Windows:
```bash
python -m venv venv
venv\Scripts\activate
```
### Instructions
```bash
pip install -r requirements.txt
```
## 4. Usage
- Open the Jupyter notebook and run all cells to reproduce the analysis:
```bash
jupyter notebook src/notebook.ipynb
```
## 5. Data
- **Source:** The Genesis portal of the German Federal Statistical Office (Destatis).
- **Content:** Employment data by year, economic sector, and employment type.
- **Format:** CSV files with well-structured tabular data.
- **Preprocessing:** Minor formatting and cleanup included in the script.
## 6. Reproducibility
- Environment can be replicated using `requirements.txt`.
## 7. Contribution
Contributions are welcome. Users can:
- Open issues
- Fork and create pull requests
## 8. License
This project is licensed under the GNU General Public License v3.0.
You can view the full license text here: [https://www.gnu.org/licenses/gpl-3.0.en.html](https://www.gnu.org/licenses/gpl-3.0.en.html)
## 9. Citation
No formal citation required for this project.
## 10. Contact
- **Name:** Joaquin Gottlebe
- **Institution:** University of Potsdam
- **Email:** joa.gottlebe@gmail.com
## 11. Acknowledgments
- Data courtesy of the German Federal Statistical Office (Destatis)
- Project developed as part of the Research Software Engineering course at the University of Potsdam
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,58 @@
*Exploratory data analysis on employment in Germany*
The dataset 13311-0001 of the Statistischen Bundesamts (Destatis)
was choosen as the dataset of this project. It was choosen because the author is currently
searching for employment and is therefore interested in insights into structures and trends
of the employment market and regional distribution. The dataset differentiates between
work-sektors and between Self-employed, employed and domestic employed.
The dataset has a temporal scope from 1991 to 2024, so 33 years.
Link: https://www-genesis.destatis.de/datenbank/online/statistic/13311/table/13311-0001
## Question 1: Which working sectors in Germany had the most and the fewest employees over the years?
**Answer**:
A **line plot** showing the number of employees per sector from 1999 to 2024 using total employment data.
- **x-axis**: Year
- **y-axis**: Number of employees
- **Color**: One line per sector
- **Tools**: `pandas`, `matplotlib`
- **Function used**: `biggest_employment_sectors_timeline()`
## Question 2: Which sectors in Germany employed the most and the fewest domestic employees over the years?
**Answer**:
Same line plot as Question 1, but filtered by **domestic concept** employment data.
- **Data filtered with**: `etype='Employees (domestic concept)'`
- **Tools**: `get_df`, `line_plot_multiple_df`
## Question 3: Which sectors had the most and the fewest self-employed individuals over the years?
**Answer**:
Line plot using self-employment data across sectors from 1999 to 2024.
- **Employment type**: `Self-employed persons a. family workers (domestic)`
- **Tools**: `get_df`, `line_plot_multiple_df`
## Question 4: Which sectors grew or declined the most in employment over the years?
**Answer**:
A **heatmap** (to be implemented) that visualizes change in employment per sector over time.
- **x-axis**: Year
- **y-axis**: Sector
- **Color scale**: Number of employees (intensity = magnitude)
- **Tools**: `pandas`, `matplotlib`
## Question 5: How did different employment types evolve over time?
**Answer**:
A line plot comparing the number of employees vs. self-employed individuals from 1999 to 2024.
- **x-axis**: Year
- **y-axis**: Number of employees
- **Lines**: One for each employment type
- **Function**: `employment_types_timeline()`
- **Tools**: `pandas`, `matplotlib`
@@ -0,0 +1,93 @@
anyio==4.9.0
appnope==0.1.4
argon2-cffi==23.1.0
argon2-cffi-bindings==21.2.0
arrow==1.3.0
asttokens==3.0.0
async-lru==2.0.5
attrs==25.3.0
babel==2.17.0
beautifulsoup4==4.13.4
bleach==6.2.0
certifi==2025.4.26
cffi==1.17.1
charset-normalizer==3.4.2
comm==0.2.2
debugpy==1.8.14
decorator==5.2.1
defusedxml==0.7.1
executing==2.2.0
fastjsonschema==2.21.1
fqdn==1.5.1
h11==0.16.0
httpcore==1.0.9
httpx==0.28.1
idna==3.10
ipykernel==6.29.5
ipython==9.2.0
ipython_pygments_lexers==1.1.1
isoduration==20.11.0
jedi==0.19.2
Jinja2==3.1.6
json5==0.12.0
jsonpointer==3.0.0
jsonschema==4.23.0
jsonschema-specifications==2025.4.1
jupyter-events==0.12.0
jupyter-lsp==2.2.5
jupyter_client==8.6.3
jupyter_core==5.7.2
jupyter_server==2.16.0
jupyter_server_terminals==0.5.3
jupyterlab==4.4.2
jupyterlab_pygments==0.3.0
jupyterlab_server==2.27.3
MarkupSafe==3.0.2
matplotlib-inline==0.1.7
mistune==3.1.3
nbclient==0.10.2
nbconvert==7.16.6
nbformat==5.10.4
nest-asyncio==1.6.0
notebook==7.4.2
notebook_shim==0.2.4
overrides==7.7.0
packaging==25.0
pandocfilters==1.5.1
parso==0.8.4
pexpect==4.9.0
platformdirs==4.3.8
prometheus_client==0.21.1
prompt_toolkit==3.0.51
psutil==7.0.0
ptyprocess==0.7.0
pure_eval==0.2.3
pycparser==2.22
Pygments==2.19.1
python-dateutil==2.9.0.post0
python-json-logger==3.3.0
PyYAML==6.0.2
pyzmq==26.4.0
referencing==0.36.2
requests==2.32.3
rfc3339-validator==0.1.4
rfc3986-validator==0.1.1
rpds-py==0.24.0
Send2Trash==1.8.3
setuptools==80.4.0
six==1.17.0
sniffio==1.3.1
soupsieve==2.7
stack-data==0.6.3
terminado==0.18.1
tinycss2==1.4.0
tornado==6.4.2
traitlets==5.14.3
types-python-dateutil==2.9.0.20241206
typing_extensions==4.13.2
uri-template==1.3.0
urllib3==2.4.0
wcwidth==0.2.13
webcolors==24.11.1
webencodings==0.5.1
websocket-client==1.8.0
@@ -0,0 +1,399 @@
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
from typing import Dict, List, Tuple
DATA_PATH = "../data/13311-0001_en_flat.csv"
primary_sectors: List[str] = [
'Agriculture, forestry and fishing',
'Mining and quarrying',
'Electricity, gas, steam, air conditioning supply',
'Water supply,sewerage,waste management,remediation'
]
secondary_sectors: List[str] = [
'Construction',
'Manufacturing',
'Industry',
'Industry, except construction'
]
tertiary_sectors: List[str] = [
'Financial, insurance, business, real estate act.',
'Financial and insurance activities',
'Real estate activities',
'Professional, scientific and technical activities',
'Administrative and support service activities',
'Business services',
'Information and communication',
'Wholesale, retail trade, repair of motor vehicles',
'Trade, transport., storage, accom.and food service',
'Trade,transport.,storage,accom.and food serv.,inf.',
'Transportation and storage',
'Accommodation and food service activities',
'Public services, education, health',
'Public admin. and defence, compulsory social sec.',
'Education',
'Human health and social work activities',
'Public a.other services, education, health,priv.h.',
'Other service activities',
'Service activities',
'Private households',
'Arts, entertainment and recreation',
'Arts,entertainm.,recreation,o.service act.,priv.h.'
]
def get_df(sector: str = 'Total',
etype: str = 'Persons in employment (domestic concept)',
start_time: int = 1999,
end_time: int = 2024) -> pd.DataFrame:
"""
Fetches and processes a DataFrame with employment data
for a specified sector and employment type over a given time range.
Args:
sector (str): The economic sector to filter
data by (default is 'Total').
etype (str): The type of employment data to retrieve
(default is 'Persons in employment (domestic concept)').
start_time (int): The starting year of the time
range (default is 1999).
end_time (int): The ending year of the time range
(default is 2024).
Returns:
pd.DataFrame: The filtered and processed DataFrame
containing the employment data.
"""
df: pd.DataFrame = pd.read_csv(DATA_PATH, sep=";")
df = df[['2_variable_attribute_label',
'value_variable_label',
'time',
'value']]
if sector != 'All':
df = df[df['2_variable_attribute_label'] == sector]
df = df[df['value_variable_label'] == etype]
df['time'] = pd.to_numeric(df['time'], errors='coerce')
df = df[(df['time'] >= start_time) & (df['time'] <= end_time)]
if sector == 'All':
df = df.sort_values(by='2_variable_attribute_label')
else:
df = df.sort_values(by='time')
df['value'] = pd.to_numeric(df['value'], errors='coerce')
if sector == 'All':
df = df.drop(['value_variable_label', 'time'], axis=1)
df = df[df['2_variable_attribute_label'] != 'Total']
else:
df = df.drop(['2_variable_attribute_label',
'value_variable_label'], axis=1)
df = df.dropna()
return df
def get_sectors() -> np.ndarray:
"""
Get all unique sector names from the dataset, excluding 'Total'.
Returns:
np.ndarray: Array of unique sector names.
"""
df: pd.DataFrame = pd.read_csv(DATA_PATH, sep=";")
df = df[['2_variable_attribute_label']]
df = df[df['2_variable_attribute_label'] != 'Total']
return pd.unique(df.values.ravel())
def line_plot_df(df: pd.DataFrame, title: str) -> None:
"""
Create a line plot for a single DataFrame.
Args:
df (pd.DataFrame): DataFrame containing 'time' and 'value' columns.
title (str): Title of the plot.
Returns:
None
"""
plt.plot(df['time'], df['value'], label=title)
plt.title(title)
plt.xlabel("Time")
plt.ylabel("Amount")
plt.legend()
plt.show()
def line_plot_multiple_df(labeled_dfs: Dict[str, pd.DataFrame],
title: str) -> None:
"""
Create a line plot for multiple DataFrames, each with its own label.
Args:
labeled_dfs (Dict[str, pd.DataFrame]): Dictionary where keys are
labels and values are DataFrames.
title (str): Title of the plot.
Returns:
None
"""
for key, df in labeled_dfs.items():
plt.plot(df['time'], df['value'], label=key)
plt.title(title)
plt.xlabel("Time")
plt.ylabel("Amount")
plt.legend()
plt.show()
def map_sector_to_group(sector: str) -> str:
"""
Map a sector name to its corresponding economic
group (Primary, Secondary, Tertiary).
Args:
sector (str): The name of the sector.
Returns:
str: The economic group ('Primary Sector', 'Secondary Sector',
'Tertiary Sector', or 'Unknown').
"""
if sector in primary_sectors:
return 'Primary Sector'
elif sector in secondary_sectors:
return 'Secondary Sector'
elif sector in tertiary_sectors:
return 'Tertiary Sector'
else:
return 'Unknown'
def get_grouped_sector_df() -> pd.DataFrame:
"""
Create a DataFrame grouping employment values by economic sector group.
Returns:
pd.DataFrame: DataFrame with groups ('Primary Sector', etc.)
and their summed employment values.
"""
df: pd.DataFrame = get_df(sector='All',
etype='Persons in employment (domestic concept)',
start_time=1999,
end_time=2024)
df = df.rename(columns={
'2_variable_attribute_label': 'sector',
'time': 'time',
'value': 'value'})
df['Group'] = df['sector'].apply(map_sector_to_group)
df = df[df['Group'] != 'Unknown']
df_grouped = df.groupby('Group')['value'].sum().reset_index()
return df_grouped
def group_values(values: List[float],
labels: List[str],
threshold: int = 5) -> Tuple[List[float], List[str]]:
"""
Group values that are below a threshold percentage
into a single '<5%' category.
Args:
values (List[float]): List of numerical values.
labels (List[str]): Corresponding labels for each value.
threshold (int): Minimum percentage to avoid grouping. Defaults to 5%.
Returns:
Tuple[List[float], List[str]]: Grouped values
and corresponding grouped labels.
"""
total: float = sum(values)
percentages: List[float] = [(value / total) * 100 for value in values]
grouped_values: List[float] = []
grouped_labels: List[str] = []
other_total: float = 0
for value, label, percent in zip(values, labels, percentages):
if percent < threshold:
other_total += value
else:
grouped_values.append(value)
grouped_labels.append(label)
if other_total > 0:
grouped_values.append(other_total)
grouped_labels.append("<5%")
return grouped_values, grouped_labels
def sort_values(values: List[float],
labels: List[str]) -> Tuple[List[float], List[str]]:
"""
Sort values and their labels in descending order based on the values.
Args:
values (List[float]): List of numerical values.
labels (List[str]): Corresponding labels for each value.
Returns:
Tuple[List[float], List[str]]: Sorted values and labels.
"""
sorted_pairs: List[Tuple[float, str]] = sorted(
zip(values, labels), reverse=True)
sorted_values, sorted_labels = zip(*sorted_pairs)
return sorted_values, sorted_labels
def pie_plot_df(values: List[float], labels: List[str], title: str) -> None:
"""
Create a pie chart from values and labels, grouping small values.
Args:
values (List[float]): List of numerical values.
labels (List[str]): Corresponding labels.
title (str): Title of the plot.
Returns:
None
"""
grouped_values, grouped_labels = group_values(values, labels, 5)
sorted_values, sorted_labels = sort_values(grouped_values, grouped_labels)
plt.pie(x=sorted_values, labels=sorted_labels, autopct='%1.1f%%')
plt.axis('equal')
plt.title(title)
plt.show()
def pie_plot_grouped_sector(df_grouped: pd.DataFrame, title: str) -> None:
"""
Create a pie chart from a grouped sector DataFrame.
Args:
df_grouped (pd.DataFrame): DataFrame with 'Group' and 'value' columns.
title (str): Title of the plot.
Returns:
None
"""
plt.figure(figsize=(8, 8))
plt.pie(x=df_grouped['value'],
labels=df_grouped['Group'],
autopct='%1.1f%%',
startangle=140)
plt.title(title)
plt.axis('equal')
plt.show()
def persons_in_employment_timeline() -> None:
"""
Plot a timeline of total persons in employment from 1999 to 2024 in Germany.
Returns:
None
"""
df_total = get_df()
line_plot_df(df_total,
'Persons in Employment from 1999-2024 in GER')
def employment_types_timeline() -> None:
"""
Plot timelines comparing employees and self-employed persons
from 1999 to 2024 in Germany.
Returns:
None
"""
empl_attr = 'Employees (domestic concept)'
self_empl_attr = 'Self-employed persons a. family workers (domestic)'
employment_data_by_type = {
empl_attr: get_df(etype=empl_attr),
self_empl_attr: get_df(etype=self_empl_attr)
}
line_plot_multiple_df(employment_data_by_type,
'Employment types from 1999-2024 in GER')
def employed_per_sector_2024() -> None:
"""
Create a pie chart showing the distribution of employees across sectors
in Germany for the year 2022.
Returns:
None
"""
df_sectors_2022 = get_df(
sector='All',
etype='Employees (domestic concept)',
start_time=2022,
end_time=2022)
labels = df_sectors_2022['2_variable_attribute_label'].tolist()
pie_plot_df(values=df_sectors_2022['value'].to_list(
), labels=labels, title='Most Employed per sector in 2022 in GER')
def self_employed_per_sector_2024() -> None:
"""
Create a pie chart showing the distribution of self-employed persons
across sectors in Germany for the year 2022.
Returns:
None
"""
df_self_sectors_2022 = get_df(
sector='All',
etype='Self-employed persons a. family workers (domestic)',
start_time=2022,
end_time=2022)
labels = df_self_sectors_2022['2_variable_attribute_label'].tolist()
pie_plot_df(values=df_self_sectors_2022['value'].to_list(
),
labels=labels,
title='Most Self-Employed per sector in 2022 in GER')
def biggest_employment_sectors_timeline() -> None:
"""
Plot timelines of employment data for the most significant sectors
in Germany from 1999 to 2024.
Returns:
None
"""
dict_sectors = {
'Service activities': get_df(sector='Service activities'),
'Private households': get_df(sector='Private households'),
'Education': get_df(sector='Education'),
'Trade,transport.,storage,accom.and food serv.,inf.': get_df(
sector='Trade,transport.,storage,accom.and food serv.,inf.'),
'Business services': get_df(sector='Business services'),
'Public services, education, health': get_df(
sector='Public services, education, health'),
'Manufacturing': get_df(sector='Manufacturing')
}
emp_sector_title = 'Biggest sectors from 1999-2024 in GER'
line_plot_multiple_df(dict_sectors, emp_sector_title)
def economic_sectors() -> None:
"""
Create a pie chart showing the total employment distribution
by economic sector group (Primary, Secondary, Tertiary).
Returns:
None
"""
df_grouped_sector = get_grouped_sector_df()
pie_plot_grouped_sector(
df_grouped_sector, 'Employment Distribution by Economic Sector')
if __name__ == "__main__":
persons_in_employment_timeline()
employment_types_timeline()
employed_per_sector_2024()
self_employed_per_sector_2024()
biggest_employment_sectors_timeline()
economic_sectors()
File diff suppressed because one or more lines are too long