This commit is contained in:
Joaquin Gottlebe
2025-07-30 15:57:48 +02:00
parent e49fcfac46
commit a639c34cee
273 changed files with 20151 additions and 0 deletions
@@ -0,0 +1,10 @@
.DS_Store
__pycache__
Thumbs.db
venv/
results/
.snakemake/
*.swp
tmp/
log.txt
workflow/report
@@ -0,0 +1,45 @@
stages:
- test
- lint
- snakemake
test_project:
stage: test
image: python:3.12-slim
tags:
- public
before_script:
- python3 -m venv venv
- source venv/bin/activate
- pip install --upgrade pip
- pip install -r requirements.txt
script:
- pytest
lint_project:
stage: lint
image: python:3.12-slim
tags:
- public
before_script:
- python3 -m venv venv
- source venv/bin/activate
- pip install --upgrade pip
- pip install pylint
- pip install -r requirements.txt
script:
- pylint workflow/scripts tests
snakemake_project:
stage: snakemake
image: python:3.12-slim
tags:
- public
before_script:
- pip install --upgrade pip
- pip install -r requirements.txt
script:
- snakemake --cores 1 --forceall
after_script:
- rm -r results
@@ -0,0 +1,2 @@
[MASTER]
init-hook='import sys; sys.path.insert(0, "workflow/scripts")'
@@ -0,0 +1,35 @@
# .cff
cff-version: 1.2.0
title: SolarLytics
message: >-
If you use this software, please cite it using the
metadata from this file.
type: software
authors:
- given-names: 'Joaquin '
family-names: Gottlebe
email: gottlebe@uni-potsdam.de
affiliation: University of Potsdam
- given-names: Flora
family-names: Grellmann
email: grellmann@uni-potsdam.de
affiliation: University of Potsdam
- given-names: Mirjam
family-names: Rupinski
email: rupinski@uni-potsdam.de
affiliation: University of Potsdam
repository-code: 'https://gitup.uni-potsdam.de/gottlebe/solarlytics'
abstract: >-
Solarlytics estimates the theoretical energy production of
German solar parks by combining area data from
OpenStreetMap with sunshine duration from the German
Weather Service. The results are compared to actual
photovoltaic electricity feed-in data from Destatis.
keywords:
- photovoltaic data
- Solarparc
- Sunshine duration
- Python
license: MIT
version: '1.0'
date-released: '2025-07-20'
@@ -0,0 +1,134 @@
# Contributor Covenant Code of Conduct
## Our Pledge
We as members, contributors, and leaders pledge to make participation in our
community a harassment-free experience for everyone, regardless of age, body
size, visible or invisible disability, ethnicity, sex characteristics, gender
identity and expression, level of experience, education, socio-economic status,
nationality, personal appearance, race, caste, color, religion, or sexual
identity and orientation.
We pledge to act and interact in ways that contribute to an open, welcoming,
diverse, inclusive, and healthy community.
## Our Standards
Examples of behavior that contributes to a positive environment for our
community include:
* Demonstrating empathy and kindness toward other people
* Being respectful of differing opinions, viewpoints, and experiences
* Giving and gracefully accepting constructive feedback
* Accepting responsibility and apologizing to those affected by our mistakes,
and learning from the experience
* Focusing on what is best not just for us as individuals, but for the overall
community
Examples of unacceptable behavior include:
* The use of sexualized language or imagery, and sexual attention or advances of
any kind
* Trolling, insulting or derogatory comments, and personal or political attacks
* Public or private harassment
* Publishing others' private information, such as a physical or email address,
without their explicit permission
* Other conduct which could reasonably be considered inappropriate in a
professional setting
## Enforcement Responsibilities
Community leaders are responsible for clarifying and enforcing our standards of
acceptable behavior and will take appropriate and fair corrective action in
response to any behavior that they deem inappropriate, threatening, offensive,
or harmful.
Community leaders have the right and responsibility to remove, edit, or reject
comments, commits, code, wiki edits, issues, and other contributions that are
not aligned to this Code of Conduct, and will communicate reasons for moderation
decisions when appropriate.
## Scope
This Code of Conduct applies within all community spaces, and also applies when
an individual is officially representing the community in public spaces.
Examples of representing our community include using an official email address,
posting via an official social media account, or acting as an appointed
representative at an online or offline event.
## Enforcement
Instances of abusive, harassing, or otherwise unacceptable behavior may be
reported to the community leaders responsible for enforcement at
[rupinski@uni-potsdam.de](mailto:rupinski@uni-potsdam.de).
All complaints will be reviewed and investigated promptly and fairly.
All community leaders are obligated to respect the privacy and security of the
reporter of any incident.
## Enforcement Guidelines
Community leaders will follow these Community Impact Guidelines in determining
the consequences for any action they deem in violation of this Code of Conduct:
### 1. Correction
**Community Impact**: Use of inappropriate language or other behavior deemed
unprofessional or unwelcome in the community.
**Consequence**: A private, written warning from community leaders, providing
clarity around the nature of the violation and an explanation of why the
behavior was inappropriate. A public apology may be requested.
### 2. Warning
**Community Impact**: A violation through a single incident or series of
actions.
**Consequence**: A warning with consequences for continued behavior. No
interaction with the people involved, including unsolicited interaction with
those enforcing the Code of Conduct, for a specified period of time. This
includes avoiding interactions in community spaces as well as external channels
like social media. Violating these terms may lead to a temporary or permanent
ban.
### 3. Temporary Ban
**Community Impact**: A serious violation of community standards, including
sustained inappropriate behavior.
**Consequence**: A temporary ban from any sort of interaction or public
communication with the community for a specified period of time. No public or
private interaction with the people involved, including unsolicited interaction
with those enforcing the Code of Conduct, is allowed during this period.
Violating these terms may lead to a permanent ban.
### 4. Permanent Ban
**Community Impact**: Demonstrating a pattern of violation of community
standards, including sustained inappropriate behavior, harassment of an
individual, or aggression toward or disparagement of classes of individuals.
**Consequence**: A permanent ban from any sort of public interaction within the
community.
## Attribution
This Code of Conduct is adapted from the [Contributor Covenant][homepage],
version 2.1, available at
[https://www.contributor-covenant.org/version/2/1/code_of_conduct.html][v2.1].
Community Impact Guidelines were inspired by
[Mozilla's code of conduct enforcement ladder][Mozilla CoC].
For answers to common questions about this code of conduct, see the FAQ at
[https://www.contributor-covenant.org/faq][FAQ]. Translations are available at
[https://www.contributor-covenant.org/translations][translations].
[homepage]: https://www.contributor-covenant.org
[v2.1]: https://www.contributor-covenant.org/version/2/1/code_of_conduct.html
[Mozilla CoC]: https://github.com/mozilla/diversity
[FAQ]: https://www.contributor-covenant.org/faq
[translations]: https://www.contributor-covenant.org/translations
@@ -0,0 +1,74 @@
# Contributing to SolarLytics
First off, thank you for your interest in contributing to **SolarLytics**!
This project is part of the Research Software Engineering course at the University of Potsdam and aims to calculate the theoretical energy production of German solar parks using open geographic and meteorological data.
Even though the project is currently maintained by a student group, we welcome all types of contributions and feedback from the broader community.
---
## Code of Conduct
This project and everyone participating in it is governed by the
[Code of Conduct](https://gitup.uni-potsdam.de/gottlebe/solarlytics/-/blob/main/CONDUCT.md).
By participating, you are expected to uphold this code. Please report unacceptable behavior
to [Mirjam Rupinski](mailto:rupinski@uni-potsdam.de).
## I Have a Questions
If you have a question or need clarification:
- First, check existing [issues](https://gitup.uni-potsdam.de/gottlebe/solarlytics/issues) — your question may have already been answered.
- If you don’t find a helpful thread, feel free to [open a new issue](https://gitup.uni-potsdam.de/gottlebe/solarlytics/issues/new).
- Include as much context as possible (e.g., error messages, system setup, relevant versions).
We’ll get back to you as soon as we can.
## What You Can Contribute
We appreciate contributions of all kinds, including (but not limited to):
- Suggestions for improvements
- Bug reports and fixes
- Code enhancements (efficiency, readability, modularity)
- Better error handling and logging
- Improvements to documentation and docstrings
- Workflow optimization (especially in Snakemake)
- Testing infrastructure
Please feel free to open an issue or pull request if you notice something that could be improved.
## Before You Start
- **Open an issue first**:\
Please open an issue to discuss any major changes before submitting a pull request.
- **Keep your code clean and structured**:
- Follow a modular and well-organized file structure.
- Document all functions thoroughly with clear and complete docstrings.
- Catch errors gracefully and log meaningful messages.
- Design your code with testing in mind.
- Consider how your additions integrate into the existing Snakemake workflows.
An example of code that can be used for orientation can be found [here](https://gitup.uni-potsdam.de/gottlebe/solarlytics/-/wikis/Coding-Style-Example).
## Local Testing
The project provides a suite of unit tests for the various module functions — check the `tests/` folder. These tests can be run easily using:
```bash
pytest
```
When adding a new function, it is expected that an appropriate and comprehensive test is also created. This helps ensure the function remains reliable after future changes.
## Contributors
Currently, we track contributors in the *Contact* section of the [README](https://gitup.uni-potsdam.de/gottlebe/solarlytics/-/blob/main/README.md) file.
---
Thanks again for your interest and support!
— The SolarLytics Team
@@ -0,0 +1,21 @@
MIT License
Copyright (c) 2025 Joaquin Gottlebe, Flora Grellmann, Mirjam Rupinski
Permission is hereby granted, free of charge, to any person obtaining a copy
of this software and associated documentation files (the "Software"), to deal
in the Software without restriction, including without limitation the rights
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
copies of the Software, and to permit persons to whom the Software is
furnished to do so, subject to the following conditions:
The above copyright notice and this permission notice shall be included in all
copies or substantial portions of the Software.
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
SOFTWARE.
@@ -0,0 +1,193 @@
# SolarLytics
## Overview
Solarlytics is a project for the Research Software Engineering course at the University of Potsdam. It aims to calculate the theoretical energy production of German solar parks. To do so, it queries the area of solar parks in Germany from OpenStreetMap (OSM) as well as the sunshine duration from the German Weather Service (DWD). The results will be compared with electricity feed-in by photovoltaics.
This project investigates the following research questions related to the theoretical energy production of solar parks in Germany:
**1. What is the spatial distribution and total surface area of solar parks in Germany?** <br>
&nbsp;&nbsp;&nbsp;&nbsp; Using geospatial data from OpenStreetMap (OSM), we aim to identify the location and size of solar parks across the country.
<br>
**2. To what extent can theoretical solar energy production be estimated for the year 2024 by combining solar park area data with sunshine duration data?**<br>
&nbsp;&nbsp;&nbsp;&nbsp; By integrating OSM data with meteorological data from the German Weather Service (DWD), we estimate the potential photovoltaic energy output for each region.
<br>
**3. How does the estimated theoretical energy production compare to the actual electricity feed-in reported by the Federal Statistical Office?** <br>
&nbsp;&nbsp;&nbsp;&nbsp; Significant deviations between theoretical and real production values may point to inefficiencies, data limitations, or external influencing factors. <br>
&nbsp;&nbsp;&nbsp;&nbsp; If significant discrepancies are observed are potential causes for discrepancies between theoretical and actual photovoltaic energy production?
<br>
**4. What are the temporal patterns of theoretical solar energy production throughout the year?** <br>
&nbsp;&nbsp;&nbsp;&nbsp; We analyze monthly trends and investigate whether seasonal or interannual differences can be observed.<br>
&nbsp;&nbsp;&nbsp;&nbsp; How do year-to-year variations in sunshine duration impact theoretical solar energy yield?
<br>
## Activity Diagram
This is a short overview of the process with the four main components of the data processing of solarparc, sunshine duration and photovoltaic data and the subsequent analysis. More details can be found in the documentation under `docs\requirements.md`.
![Alt-Text](/docs/UML_diagram_overview.png)
<br>
## Getting Started
### Requirements
- **Python**    3.13.5
- **git**      2.39.5
All additional Python package dependencies are listed in `requirements.txt`.
### Installation
#### Clone the repositpory
The repository can be cloned with the following command:
```bash
git clone https://gitup.uni-potsdam.de/gottlebe/solarlytics.git
```
#### Change directory
```bash
cd solarlytics
```
#### Set up enviroment
The `requirements.txt` file can now be used to create an environment as follows:
Creates a virtual environment named `venv` using Python
```bash
python3 -m venv venv
```
Activates the virtual environment, so any `pip` or `python` commands now use the local environment, not the global Python installation.
```bash
source venv/bin/activate
```
Installs all Python packages listed in the `requirements.txt` file into the virtual environment.
```bash
pip3 install -r requirements.txt
```
## Usage
To run the complete analysis pipeline, simply execute:
```bash
snakemake --cores 1 --forceall
```
This runs Snakemake, using 1 CPU core, and forces all steps to run again, even if their outputs already exist.
This will automatically do the calculation described in the UML-Diagram:
- Download and process the required datasets (OpenStreetMap and DWD),
- Compute the theoretical solar energy production for German solar parks,
- Compare results with actual photovoltaic electricity feed-in data,
- Generate plots, diagrams, and a final report.
All new data frames are saved in the `results/` directory. The report's images and files are saved under `workflow/reports/` and can also be viewed in the automatically generated *slides* (.md) under `docs/`.
## Specific Usage
To remove all results, run:
```bash
snakemake --cores 1 clean
```
To run all unittests, use the following command:
```bash
pytest
```
### Parameterisation
You can customize the analysis by modifying parameters in the `config/config.yml` file. Specifically:
**year**: Defines the year to be analyzed.
**efficiency**: Specifies the average efficiency of the solar parcs, energy feed in and solar radiation (defined as a decimal, e.g. 0.15 for 15%).
Adjusting these values allows flexible evaluation for different scenarios or datasets.
## Data
This Project uses three datasets:
#### Solarparc Data
To get the area of all solarparks in germany, we use data from Open Street Maps. We use polygons (outlines) of the solarparcs and the outline of germany.
| Data | Solarparc Data |
|----------|----------|
| Origin | Open Steet Maps |
| Data format | OSM XML |
| API | OverpassAPI|
| Link | https://www.openstreetmap.org/ |
| Licence | Data © OpenStreetMap contributors, licensed under Open Database License (ODbL) v1.0 |
#### Photovoltaik Electricity Feed-in Data
Statistics on the monthly electricity feed-in of various energy sources by the Federal Statistical Office of Germany (Destatis). It is possible to show or hide different properties. In this project, the electricity feed-in from photovoltaic systems monthly and over several years is of interest.
| Data | Photovoltaic Data |
|----------|----------|
| Origin | DESTATIS |
| Data format | CSV oder XML |
| Link | https://www-genesis.destatis.de/datenbank/online/statistic/43312/table/43312-0001 |
| Licence | Data Licence Germany 2.0. |
#### Sunshine Duration Data
The DWD provides many different datasets on sunshine duration. We decided to use the monthly German averages.
| Data | Sunshine Duration Data |
|----------|----------|
| Origin | DWD |
| Data format | TXT |
| Link | [opendata.dwd.de/climate_environment/CDC/regional_averages_DE/monthly/sunshine_duration/](opendata.dwd.de/climate_environment/CDC/regional_averages_DE/monthly/sunshine_duration/) |
| Licence | CC-BY-4.0 |
## Contribution
Contributions are welcome and encouraged! If you would like to contribute to this project, please first read our [Contribution Guidelines](CONTRIBUTING.md) to understand the development workflow and code standards.
We also expect all contributors to follow our [Code of Conduct](CONDUCT.md) to ensure a respectful and inclusive environment.
If you're looking for a place to start, check out the [open issues](https://gitup.uni-potsdam.de/gottlebe/solarlytics/-/issues).
## License
**OpenStreetMap data**:
Data © OpenStreetMap contributors, licensed under the Open Database License (ODbL) v1.0.
See https://www.openstreetmap.org/copyright for details.
**Project code and reports**:
This project is licensed under the MIT License – see [LICENSE](https://gitup.uni-potsdam.de/gottlebe/solarlytics/-/blob/main/LICENSE) file in this repository.
This means you can freely use and modify the code.
If you ever publish processed OSM datasets, those must remain under ODbL with proper attribution.
## Citation
If you use this software, please cite it as described in the [CITATION.cff](https://gitup.uni-potsdam.de/gottlebe/solarlytics/-/blob/tree/main/CITATION.cff) file.
Example citation (in APA format):
Gottlebe, J. Grellmann F., Rupinski M. (2025). SolarLytics: Estimating theoretical solar energy production in Germany (Version 1.0) [Computer software]. https://gitup.uni-potsdam.de/gottlebe/solarlytics
## Contact
For questions or feedback, feel free to reach out via e-mail:
#### Owner:
- Joaquin Gottlebe: gottlebe@uni-potsdam.de
#### Maintainer:
- Flora Grellmann: grellmann@uni-potsdam.de
- Mirjam Rupinski: rupinski@uni-potsdam.de
#### Contributors:
\-
@@ -0,0 +1,7 @@
result_path: "results"
report_path: "workflow/report"
data_path: "data"
scripts_path: "workflow/scripts"
year: "2024"
efficiency: 0.00216
Binary file not shown.

After

Width:  |  Height:  |  Size: 126 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 49 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 82 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 39 KiB

@@ -0,0 +1,68 @@
# SolarLytics Requirements
## Functional Requirements
#### UML Diagram
Main diagram describing the process of the complete project:
<br> ![Alt-Text](/docs/UML_diagram_overview.png)
The individual processing steps of the different datasets are shown here in more detail:
<br> ![Alt-Text](/docs/UML_diagram_detail.png)
## Non-functional Requirements
### Must Have
- **Compatibility**: The workflow must run on Windows and Mac OS. It should be easy to deploy and run in different environments.
- **Documentation**: The documentation is for users and developers and must be available, including metadata, license information, a read-me file, and citation as well as contributing guidelines.
- **Reproducibility**: The workflow must be reproducible.
- **Testability**: The code must be well tested, with unit and integration tests in place.
### Should Have
- **Performance**: The workflow should able to process the datasets in a reasonable time.
- **Maintainability**: The code should be well modularized and organized, following the standard template and PEP 8 guidelines. It should be refactured regularily.
- **Easy-to-Use**: The workflow should provide a easy user experience
### Could Have
- **Accessibility**: The documentation could be translated to German.
- **Presentation**: The workflow could have a visual support in the form of graphics.
### Won't Have
- **Add-ons**: Function for other countries.
# Component Analysis
| Abstract Workflow Node (Operation) | Input(s) | Output(s) |Implementation |
|------------------------------------|-------------------------------------------|---------------------------------------|----------------------------------------------|
|**photovoltaic data processing** | | |
| load and clean photovoltaics data | original dataset (.csv) from destatis | dataframe (.csv) | CLI tool built on pandas
| trim photovoltaic | dataframe (.csv) | dataframe(.csv) | CLI tool built on pandas
| collapse columns photovoltaic | dataframe (.csv) | dataframe (.csv) | CLI tool built on pandas
| | | |
|**sunshine duration data processing** | | |
| load sunshine duration data | DWD Website | sunshine duration data (.txt), one file for each month | CLI tool built on requests, os, bs4 and urllib
| merge series | 12 sunshine duration datasets (.txt) | processed dataframe (.csv) | CLI tool built on pandas
| collapse columns sunshine duration | dataframe (.csv) | dataframe (.csv) | CLI tool built on pandas
| trim sunshine duration | dataframe (.csv) | dataframe (.csv) | CLI tool built on pandas
| | | |
|**solarparc data processing** | | |
| load solarparc data and border data| OverpassAPI | .gpkg file (solarparc / border) | CLI http request over OverpassAPI, osmium, geopandas, shapely
| calculate area | .gpkg file | .txt with report | CLI tool built on geopandas
| | | |
|**analysis** | | |
| calculate theoretical energy | calculate area (.txt) & sunshine duration dataframe (.csv) | dataframe (.csv) | CLI tool built on pandas
| collapse columns theoretical energy| dataframe (.csv) | dataframe (.csv) | CLI tool built on pandas
| trim theoreticalenergy | dataframe (.csv) | energy differnce dataframe (.csv) | CLI tool built in pandas
| calculate difference | theoritcal energy (.csv) & photovoltaic energy (.csv) | dataframe .csv | CLI tool built on pandas
| plot energy difference | ernergy difference (.csv) | diagram over years (.png) | CLI tool built on matplotlib and pandas
| plot yearly energy change | dataframe (.csv) | diagram over years (.png) | CLI tool built on matplotlib and pandas
| plot solarparc map | .gpkg file (solarparc / border) | plot of germany (.png) | CLI tool built on matplotlib and geopandas
| plot yearly change sunshine | dataframe (.csv) | diagram over years (.png) | CLI tool built on matplotlib and pandas
| plot yearly change photovoltaic | dataframe (.csv) | diagram over years (.png) | CLI tool built on matplotlib and pandas
| plot monthly change potovoltaic | dataframe (.csv) | diagram over months (.png) | CLI tool built on matplotlib and pandas
| generate report | mutiple inputs (.txt, .png) | slides (.md) | built on markdown
@@ -0,0 +1,169 @@
---
marp: true
theme: uncover
class: invert
paginate: true
_paginate: false
header:
footer: '14.07.2025 / University of Potsdam / Flora Grellmann, Joaquin Gottlebe, Mirjam Rupinski'
---
<!-- _paginate: skip -->
# Germany’s Solar Potential Theory vs. Reality
---
##### Research Questions
1. To what extent can theoretical solar energy production be estimated?
2. How does the theoretical energy production compare to the actual energy production?
...
---
##### Modules
![w:490](../docs/UML_diagram_overview.png)
---
##### Non functional requirements
**Compatibility**: The workflow must run on Windows Mac OS and Linux.
**Documentation**: The documentation is for users and developers.
**Reproducibility**: The workflow must be reproducible.
**Testability**: The code must be well tested.
---
##### Workflow: Part 1
![w:1190](../docs/UML_diagram_detail_1.png)
---
##### Workflow: Part 2
![w:580](../docs/UML_diagram_detail_2.png)
---
##### Dataset
<style scoped>
section {
font-size: 20px;
}
</style>
| Dataset | Origin | Data format | API | Link | Licence |
| -------------------- | ---------------- | ----------- | ----------- | ------------------- | ---------------------------- |
| Solarparcs & Germany | Open Street Maps | OSM XML | OverpassAPI | openstreetmap.org | © OpenStreetMap contributors |
| Sunshine Duration | DWD | TXT | - | opendata.dwd.de | CC-BY-4.0 |
| Actual Energy | DESTATIS | CSV or XML | - | genesis.destatis.de | Data Licence Germany 2.0 |
---
<style scoped>
section {
font-size: 20px;
}
</style>
#### Created Tools
| **Stage** | **Tool** |
| ------------------------- | ----------------------- |
| **Data Acquisition** | wgetdir.py |
| **Data Processing** | clip.py |
| | merge_series.py |
| | clean_pv_data.py |
| | polygons2area.py |
| | collapse_columns.py |
| | trim.py |
| | calculate_energy.py |
| | calculate_difference.py |
| **Data Visualization** | plot_geo.py |
| | plot_yearly_change.py |
| | plot_monthly_change.py |
| | plot_difference.py |
---
<!-- Flo -->
##### Results - Solarparcs Area
- Total area: 16.000 $km^2$
![bg right](../workflow/report/solarparcs.png)
---
##### Results - Sunshine Duration
![](../workflow/report/sunshine_duration_yearly_trimmed.png)
---
##### Results - Theoretical Energy
![](../workflow/report/solarparcs_energy_yearly_trimmed.png)
---
##### Results - Actual Energy
![](../workflow/report/pv_data_cleaned_trimmed_collapsed.png)
---
##### Results - Difference
![](../workflow/report/energy_difference.png)
---
<style scoped>
section {
font-size: 30px;
}
</style>
##### Reflection
###### Worked well
- Working with workflows (Snakemake)
- Working with Issues and Tasks as a group
- Review process
- Access to relevant data
###### Challenges
- Defining clear roadmap
- Group communication
- Working with licences
---
##### Not included
- Analysis of monthly change
---
##### Thank you for you attention.
Why do you think the theorical and actual energy diverges more in earlier year?
@@ -0,0 +1,7 @@
[pytest]
pythonpath = workflow/scripts
filterwarnings = ignore::DeprecationWarning:pyogrio.*
log_cli = true
log_cli_level = INFO
@@ -0,0 +1,7 @@
pytest==8.4.1
geopandas==1.1.0
matplotlib==3.10.3
bs4==0.0.2
snakemake==9.6.1
logging==0.4.9.6
pylint==3.3.7
@@ -0,0 +1,27 @@
"""
Shared test fixtures for pytest.
This file defines reusable fixtures that are automatically discovered by pytest
and can be used across multiple test modules.
Fixtures:
---------
- local_tmp_path: Creates (or cleans) a temporary directory under 'tests/tmp'
to store output files generated during tests. Ensures a clean environment
for each test run.
"""
from pathlib import Path
import pytest
@pytest.fixture
def local_tmp_path():
"""Fixture to create and return a temporary path for output files."""
path = Path("tests/tmp")
if path.exists():
for file in path.iterdir():
file.unlink()
else:
path.mkdir(parents=True)
return path
@@ -0,0 +1 @@
197429970.81
@@ -0,0 +1 @@
1
@@ -0,0 +1,4 @@
Jahr,January,February,March,April,May,June,July,August,September,October,November,December
1951,3.09,5.73,11.01,19.91,20.33,21.02,23.88,20.41,17.51,19.25,3.85,5.18
1952,3.30,5.20,12.43,19.57,21.20,22.47,26.14,20.75,11.74,7.96,3.79,3.12
1 Jahr January February March April May June July August September October November December
2 1951 3.09 5.73 11.01 19.91 20.33 21.02 23.88 20.41 17.51 19.25 3.85 5.18
3 1952 3.30 5.20 12.43 19.57 21.20 22.47 26.14 20.75 11.74 7.96 3.79 3.12
@@ -0,0 +1,3 @@
Jahr,January,February,March,April,May,June,July,August,September,October,November,December
1951,30.9,57.3,110.1,199.1,203.3,210.2,238.8,204.1,175.1,192.5,38.5,51.8
1952,33.0,52.0,124.3,195.7,212.0,224.7,261.4,207.5,117.4,79.6,37.9,31.2
1 Jahr January February March April May June July August September October November December
2 1951 30.9 57.3 110.1 199.1 203.3 210.2 238.8 204.1 175.1 192.5 38.5 51.8
3 1952 33.0 52.0 124.3 195.7 212.0 224.7 261.4 207.5 117.4 79.6 37.9 31.2
@@ -0,0 +1,7 @@
Zeitreihen fuer Gebietsmittel fuer Bundeslaender und Kombinationen von Bundeslaender, erstellt am: 20250602
Jahr;Monat;Brandenburg/Berlin;Brandenburg;Baden-Wuerttemberg;Bayern;Hessen;Mecklenburg-Vorpommern;Niedersachsen;Niedersachsen/Hamburg/Bremen;Nordrhein-Westfalen;Rheinland-Pfalz;Schleswig-Holstein;Saarland;Sachsen;Sachsen-Anhalt;Thueringen/Sachsen-Anhalt;Thueringen;Deutschland;
1951;01; 45.2; 45.3; 30.6; 34.5; 22.8; 39.7; 24.4; 24.6; 23.3; 23.9; 35.8; 21.5; 37.6; 28.0; 26.0; 23.5; 30.9;
1952;01; 38.9; 39.0; 39.1; 39.3; 19.3; 38.1; 25.9; 26.0; 23.3; 27.0; 38.1; 15.2; 41.9; 35.0; 30.7; 25.3; 33.0;
@@ -0,0 +1,7 @@
Zeitreihen fuer Gebietsmittel fuer Bundeslaender und Kombinationen von Bundeslaender, erstellt am: 20250602
Jahr;Monat;Brandenburg/Berlin;Brandenburg;Baden-Wuerttemberg;Bayern;Hessen;Mecklenburg-Vorpommern;Niedersachsen;Niedersachsen/Hamburg/Bremen;Nordrhein-Westfalen;Rheinland-Pfalz;Schleswig-Holstein;Saarland;Sachsen;Sachsen-Anhalt;Thueringen/Sachsen-Anhalt;Thueringen;Deutschland;
1951;02; 67.0; 67.2; 65.9; 63.7; 49.7; 56.1; 44.6; 44.5; 52.1; 59.1; 37.5; 59.8; 80.6; 50.2; 50.5; 50.9; 57.3;
1952;02; 61.5; 61.5; 61.1; 42.5; 43.6; 62.9; 52.3; 52.4; 42.9; 57.4; 65.0; 74.9; 46.4; 58.5; 52.0; 43.9; 52.0;
@@ -0,0 +1,8 @@
Zeitreihen fuer Gebietsmittel fuer Bundeslaender und Kombinationen von Bundeslaender, erstellt am: 20250602
Jahr;Monat;Brandenburg/Berlin;Brandenburg;Baden-Wuerttemberg;Bayern;Hessen;Mecklenburg-Vorpommern;Niedersachsen;Niedersachsen/Hamburg/Bremen;Nordrhein-Westfalen;Rheinland-Pfalz;Schleswig-Holstein;Saarland;Sachsen;Sachsen-Anhalt;Thueringen/Sachsen-Anhalt;Thueringen;Deutschland;
1951;03; 129.5; 129.5; 96.3; 98.9; 114.6; 117.4; 114.6; 114.9; 105.0; 107.3; 117.9; 110.3; 116.9; 118.1; 115.7; 112.7; 110.1;
1952;03; 163.3; 163.3; 115.7; 107.0; 110.7; 178.7; 121.5; 121.9; 99.7; 106.1; 152.4; 111.7; 137.6; 136.8; 125.3; 110.9; 124.3;
@@ -0,0 +1,8 @@
Zeitreihen fuer Gebietsmittel fuer Bundeslaender und Kombinationen von Bundeslaender, erstellt am: 20250602
Jahr;Monat;Brandenburg/Berlin;Brandenburg;Baden-Wuerttemberg;Bayern;Hessen;Mecklenburg-Vorpommern;Niedersachsen;Niedersachsen/Hamburg/Bremen;Nordrhein-Westfalen;Rheinland-Pfalz;Schleswig-Holstein;Saarland;Sachsen;Sachsen-Anhalt;Thueringen/Sachsen-Anhalt;Thueringen;Deutschland;
1951;04; 227.3; 227.4; 195.6; 198.0; 189.0; 226.3; 185.3; 185.4; 180.3; 187.5; 189.3; 196.8; 223.6; 208.2; 204.0; 198.7; 199.1;
1952;04; 204.4; 204.6; 180.2; 181.8; 192.0; 207.6; 205.3; 205.5; 203.2; 199.5; 209.7; 223.8; 195.2; 200.8; 195.4; 188.6; 195.7;
@@ -0,0 +1,8 @@
Zeitreihen fuer Gebietsmittel fuer Bundeslaender und Kombinationen von Bundeslaender, erstellt am: 20250602
Jahr;Monat;Brandenburg/Berlin;Brandenburg;Baden-Wuerttemberg;Bayern;Hessen;Mecklenburg-Vorpommern;Niedersachsen;Niedersachsen/Hamburg/Bremen;Nordrhein-Westfalen;Rheinland-Pfalz;Schleswig-Holstein;Saarland;Sachsen;Sachsen-Anhalt;Thueringen/Sachsen-Anhalt;Thueringen;Deutschland;
1951;05; 195.3; 195.5; 197.8; 192.2; 202.7; 244.1; 211.1; 211.5; 203.8; 205.3; 243.9; 216.2; 200.5; 180.5; 182.2; 184.4; 203.3;
1952;05; 203.6; 203.8; 226.9; 196.6; 222.7; 214.4; 227.7; 227.9; 211.8; 221.0; 237.9; 247.1; 182.4; 204.0; 202.4; 200.4; 212.0;
@@ -0,0 +1,8 @@
Zeitreihen fuer Gebietsmittel fuer Bundeslaender und Kombinationen von Bundeslaender, erstellt am: 20250602
Jahr;Monat;Brandenburg/Berlin;Brandenburg;Baden-Wuerttemberg;Bayern;Hessen;Mecklenburg-Vorpommern;Niedersachsen;Niedersachsen/Hamburg/Bremen;Nordrhein-Westfalen;Rheinland-Pfalz;Schleswig-Holstein;Saarland;Sachsen;Sachsen-Anhalt;Thueringen/Sachsen-Anhalt;Thueringen;Deutschland;
1951;06; 245.1; 245.3; 194.2; 198.7; 190.0; 255.5; 212.8; 213.4; 191.4; 190.1; 251.0; 190.8; 223.8; 210.6; 203.7; 195.0; 210.2;
1952;06; 233.1; 233.0; 254.6; 227.7; 226.9; 223.7; 202.2; 202.2; 204.6; 236.4; 198.5; 265.3; 230.9; 233.8; 229.2; 223.3; 224.7;
@@ -0,0 +1,8 @@
Zeitreihen fuer Gebietsmittel fuer Bundeslaender und Kombinationen von Bundeslaender, erstellt am: 20250602
Jahr;Monat;Brandenburg/Berlin;Brandenburg;Baden-Wuerttemberg;Bayern;Hessen;Mecklenburg-Vorpommern;Niedersachsen;Niedersachsen/Hamburg/Bremen;Nordrhein-Westfalen;Rheinland-Pfalz;Schleswig-Holstein;Saarland;Sachsen;Sachsen-Anhalt;Thueringen/Sachsen-Anhalt;Thueringen;Deutschland;
1951;07; 252.6; 252.6; 270.6; 258.0; 238.3; 248.3; 205.1; 205.3; 202.2; 241.8; 213.4; 270.7; 248.7; 229.5; 233.3; 238.1; 238.8;
1952;07; 269.1; 269.2; 306.9; 296.0; 267.1; 259.3; 209.4; 209.4; 211.0; 266.1; 236.3; 302.3; 281.3; 245.1; 251.7; 259.9; 261.4;
@@ -0,0 +1,8 @@
Zeitreihen fuer Gebietsmittel fuer Bundeslaender und Kombinationen von Bundeslaender, erstellt am: 20250602
Jahr;Monat;Brandenburg/Berlin;Brandenburg;Baden-Wuerttemberg;Bayern;Hessen;Mecklenburg-Vorpommern;Niedersachsen;Niedersachsen/Hamburg/Bremen;Nordrhein-Westfalen;Rheinland-Pfalz;Schleswig-Holstein;Saarland;Sachsen;Sachsen-Anhalt;Thueringen/Sachsen-Anhalt;Thueringen;Deutschland;
1951;08; 213.9; 214.1; 213.8; 220.7; 201.5; 203.9; 192.5; 192.8; 179.8; 198.1; 196.3; 213.4; 226.3; 186.6; 188.6; 191.2; 204.1;
1952;08; 220.3; 220.1; 238.0; 227.1; 195.1; 214.9; 184.4; 184.7; 176.2; 198.4; 186.7; 218.7; 223.0; 196.2; 195.8; 195.3; 207.5;
@@ -0,0 +1,8 @@
Zeitreihen fuer Gebietsmittel fuer Bundeslaender und Kombinationen von Bundeslaender, erstellt am: 20250602
Jahr;Monat;Brandenburg/Berlin;Brandenburg;Baden-Wuerttemberg;Bayern;Hessen;Mecklenburg-Vorpommern;Niedersachsen;Niedersachsen/Hamburg/Bremen;Nordrhein-Westfalen;Rheinland-Pfalz;Schleswig-Holstein;Saarland;Sachsen;Sachsen-Anhalt;Thueringen/Sachsen-Anhalt;Thueringen;Deutschland;
1951;09; 205.3; 205.3; 170.7; 173.1; 165.8; 196.2; 167.8; 168.0; 149.1; 173.7; 166.9; 190.9; 182.6; 191.9; 183.2; 172.3; 175.1;
1952;09; 136.3; 136.3; 116.1; 101.1; 93.5; 177.3; 128.1; 128.7; 103.2; 106.0; 161.1; 117.8; 103.5; 116.3; 103.3; 87.0; 117.4;
@@ -0,0 +1,8 @@
Zeitreihen fuer Gebietsmittel fuer Bundeslaender und Kombinationen von Bundeslaender, erstellt am: 20250602
Jahr;Monat;Brandenburg/Berlin;Brandenburg;Baden-Wuerttemberg;Bayern;Hessen;Mecklenburg-Vorpommern;Niedersachsen;Niedersachsen/Hamburg/Bremen;Nordrhein-Westfalen;Rheinland-Pfalz;Schleswig-Holstein;Saarland;Sachsen;Sachsen-Anhalt;Thueringen/Sachsen-Anhalt;Thueringen;Deutschland;
1951;10; 227.4; 227.5; 151.6; 183.7; 184.3; 225.7; 204.9; 205.4; 181.9; 170.1; 215.6; 169.7; 206.0; 200.4; 195.0; 188.2; 192.5;
1952;10; 55.8; 55.9; 94.2; 89.6; 81.9; 76.2; 71.9; 71.8; 81.0; 100.1; 66.5; 103.2; 71.0; 72.9; 72.9; 72.9; 79.6;
@@ -0,0 +1,8 @@
Zeitreihen fuer Gebietsmittel fuer Bundeslaender und Kombinationen von Bundeslaender, erstellt am: 20250602
Jahr;Monat;Brandenburg/Berlin;Brandenburg;Baden-Wuerttemberg;Bayern;Hessen;Mecklenburg-Vorpommern;Niedersachsen;Niedersachsen/Hamburg/Bremen;Nordrhein-Westfalen;Rheinland-Pfalz;Schleswig-Holstein;Saarland;Sachsen;Sachsen-Anhalt;Thueringen/Sachsen-Anhalt;Thueringen;Deutschland;
1951;11; 32.7; 32.8; 54.7; 48.5; 27.5; 37.7; 28.7; 28.7; 39.2; 44.3; 28.3; 45.9; 35.0; 29.6; 29.0; 28.3; 38.5;
1952;11; 31.3; 31.4; 44.6; 40.8; 34.1; 31.4; 32.2; 32.2; 43.4; 38.1; 39.6; 32.3; 44.0; 33.4; 36.1; 39.4; 37.9;
@@ -0,0 +1,8 @@
Zeitreihen fuer Gebietsmittel fuer Bundeslaender und Kombinationen von Bundeslaender, erstellt am: 20250602
Jahr;Monat;Brandenburg/Berlin;Brandenburg;Baden-Wuerttemberg;Bayern;Hessen;Mecklenburg-Vorpommern;Niedersachsen;Niedersachsen/Hamburg/Bremen;Nordrhein-Westfalen;Rheinland-Pfalz;Schleswig-Holstein;Saarland;Sachsen;Sachsen-Anhalt;Thueringen/Sachsen-Anhalt;Thueringen;Deutschland;
1951;12; 51.2; 51.4; 69.9; 42.7; 38.1; 53.2; 51.6; 51.5; 59.5; 49.6; 42.7; 65.1; 61.3; 52.3; 50.7; 48.7; 51.8;
1952;12; 35.8; 35.9; 34.9; 28.4; 30.6; 29.9; 24.0; 24.0; 32.0; 33.7; 21.7; 33.7; 50.5; 29.8; 31.6; 33.9; 31.2;
@@ -0,0 +1,83 @@
"""
Unit tests for the calculate_difference.py module, which verifies the correctness
of the energy difference calculation.
"""
from unittest import mock
import pandas as pd
from calculate_difference import calc_energy_difference
@mock.patch("calculate_difference.logs.log_processed")
@mock.patch("calculate_difference.logs.log_processing")
@mock.patch("calculate_difference.checks.check_dir")
@mock.patch("calculate_difference.checks.check_path")
@mock.patch("calculate_difference.utils.save_df")
@mock.patch("calculate_difference.utils.read_df")
def test_calc_energy_difference(mock_read_df, mock_save_df,
_mock_check_path, _mock_check_dir,
_mock_log_processing, _mock_log_processed):
"""
Test that calc_difference_difference computes the correct energy difference per
year and saves the expected DataFrame.
"""
df_actual = pd.DataFrame({
'year': [2018, 2019, 2020],
'Summe': [100, 200, 300]
})
df_theoretical = pd.DataFrame({
'year': [2018, 2019, 2020],
'Summe': [90, 210, 310]
})
mock_read_df.side_effect = [df_actual, df_theoretical]
calc_energy_difference("actual.csv", "theoretical.csv", "output.csv")
expected_diff = abs(df_actual['Summe'] - df_theoretical['Summe'])
expected_df = pd.DataFrame({
'year': df_actual["year"],
'Energy Difference [kWh]': expected_diff
})
result_df = mock_save_df.call_args[0][0]
# adjust types, because "2018" =/= 2018
expected_df["year"] = expected_df["year"].astype(int)
result_df["year"] = result_df["year"].astype(int)
pd.testing.assert_frame_equal(
result_df.sort_values(by="year").reset_index(drop=True),
expected_df.sort_values(by="year").reset_index(drop=True)
)
@mock.patch("calculate_difference.logs.log_processed")
@mock.patch("calculate_difference.logs.log_processing")
@mock.patch("calculate_difference.checks.check_dir")
@mock.patch("calculate_difference.checks.check_path")
@mock.patch("calculate_difference.utils.save_df")
@mock.patch("calculate_difference.utils.read_df")
def test_calc_energy_difference_dimensions(mock_read_df, mock_save_df,
_mock_check_path, _mock_check_dir,
_mock_log_processing, _mock_log_processed):
"""
Test that the resulting DataFrame has the expected shape: same row count as input
and two columns ('year' and 'Energy Difference [kWh]').
"""
df_actual = pd.DataFrame({
'year': [2018, 2019, 2020],
'Summe': [100, 200, 300]
})
df_theoretical = pd.DataFrame({
'year': [2018, 2019, 2020],
'Summe': [90, 210, 310]
})
mock_read_df.side_effect = [df_actual, df_theoretical]
calc_energy_difference("actual.csv", "theoretical.csv", "output.csv")
result_df = mock_save_df.call_args[0][0]
# checks whether the number of rows is the same
assert len(result_df) == len(df_actual)
# checks exactly 2 columns ("year", "energy difference")
assert result_df.shape[1] == 2
@@ -0,0 +1,23 @@
"""Unit tests for the calculate_energy function """
import os
import pandas as pd
from calculate_energy import calculate_energy
def test_calculate_energy(local_tmp_path):
"""
Test that calculate_energy processes solarparc area and sunshine data to compute
energy output, and writes the result to a non-empty CSV file.
"""
area_data_path = "tests/data/area_easy.txt"
sunshine_data_path = "tests/data/sunshine_duration.csv"
output_path = local_tmp_path / "energy.csv"
expected_path = "tests/data/solarparc_energy.csv"
calculate_energy(area_data_path, sunshine_data_path, output_path, 1)
expected_data = pd.read_csv(expected_path)
output_data = pd.read_csv(output_path)
assert os.path.exists(output_path)
pd.testing.assert_frame_equal(expected_data, output_data)
@@ -0,0 +1,160 @@
"""
Unit tests for the `checks` module, which validate various input verification functions.
"""
import sys
import pandas as pd
import geopandas as gpd
from shapely.geometry import Point
import pytest
import checks
def test_check_empty():
"""
Test that check_empty raises ValueError for empty or None inputs, and
passes for valid inputs.
"""
checks.check_empty("hello")
checks.check_empty([1, 2, 3])
try:
checks.check_empty("")
except ValueError:
pass
else:
assert False, "Expected ValueError for empty string"
try:
checks.check_empty([])
except ValueError:
pass
else:
assert False, "Expected ValueError for empty list"
try:
checks.check_empty(None)
except ValueError:
pass
else:
assert False, "Expected ValueError for None"
def test_check_path(local_tmp_path):
"""
Test that check_path passes for existing files and raises FileNotFoundError
for missing files.
"""
tmp_file = local_tmp_path / "file.txt"
tmp_file.write_text("test")
checks.check_path(str(tmp_file))
with pytest.raises(FileNotFoundError):
checks.check_path(str(local_tmp_path / "nonexistent.txt"))
def test_check_dir(local_tmp_path):
"""Test that check_dir verifies if the directory part of a path exists."""
some_file = local_tmp_path / "some_file.txt"
checks.check_dir(str(some_file))
with pytest.raises(FileNotFoundError):
checks.check_dir("/non/existing/parent_dir/file.txt")
def test_check_args():
"""Test that check_args validates the number of command-line arguments."""
original_argv = sys.argv
sys.argv = ['prog', 'arg1', 'arg2']
checks.check_args(3, "Wrong args count")
try:
checks.check_args(2, "Wrong args count")
except RuntimeError:
pass
else:
assert False, "Expected RuntimeError for wrong arg count"
sys.argv = original_argv
def test_check_crs():
"""
Test that check_crs passes for matching CRS values and raises ValueError
on mismatch.
"""
checks.check_crs("EPSG:4326", "EPSG:4326")
try:
checks.check_crs("EPSG:4326", "EPSG:3857")
except ValueError:
pass
else:
assert False, "Expected ValueError for CRS mismatch"
def test_check_df():
"""Test that check_df validates input as non-empty pandas DataFrame with non-NaN data."""
df = pd.DataFrame({"a": [1, 2]})
checks.check_df(df)
try:
checks.check_df([1, 2, 3])
except TypeError:
pass
else:
assert False, "Expected TypeError for non-DataFrame"
try:
checks.check_df(pd.DataFrame())
except ValueError:
pass
else:
assert False, "Expected ValueError for empty DataFrame"
try:
checks.check_df(pd.DataFrame({"a": [None, None]}))
except ValueError:
pass
else:
assert False, "Expected ValueError for DataFrame all NaNs"
def test_check_gdf():
"""Test that check_gdf validates input as non-empty GeoDataFrame with valid geometries."""
gdf = gpd.GeoDataFrame({'geometry': [Point(0, 0)]})
checks.check_gdf(gdf)
try:
checks.check_gdf(pd.DataFrame())
except TypeError:
pass
else:
assert False, "Expected TypeError for non-GeoDataFrame"
try:
checks.check_gdf(gpd.GeoDataFrame())
except ValueError:
pass
else:
assert False, "Expected ValueError for empty GeoDataFrame"
empty_geom = gpd.GeoDataFrame({'geometry': [None, None]})
try:
checks.check_gdf(empty_geom)
except ValueError:
pass
else:
assert False, "Expected ValueError for GeoDataFrame with empty geometries"
def test_check_member():
"""Test that check_member confirms presence of a member in a collection."""
checks.check_member('a', ['a', 'b', 'c'])
try:
checks.check_member('d', ['a', 'b', 'c'])
except ValueError:
pass
else:
assert False, "Expected ValueError for invalid member"
@@ -0,0 +1,45 @@
"""
Unit tests for the module clean_pv_data.py
"""
from unittest.mock import patch
import pandas as pd
from clean_pv_data import clean_pv_data
@patch("utils.save_df")
@patch("utils.read_df")
@patch("checks.check_empty")
@patch("checks.check_dir")
@patch("checks.check_path")
@patch("logs.log_processing")
@patch("logs.log_processed")
def test_clean_pv_data(_mock_log_processed, _mock_log_processing,
_mock_check_path, _mock_check_dir, _mock_check_empty,
mock_read_df, mock_save_df):
"""
Test that clean_pv_data filters by variable label and reshapes the data into
a wide format by month.
"""
# Dummy parameter setup
extracted_type = "Electricity feed-in"
inputh_path = "dummy_input.csv"
output_path = "dummy_output.csv"
input_data = pd.DataFrame({
"time": ["2022", "2022", "2022"],
"1_variable_attribute_label": ["January", "February", "January"],
"value": [100, 150, 300],
"value_variable_label": ["Electricity feed-in",
"Electricity feed-in",
"Net nominal capacity"]
})
mock_read_df.return_value = input_data
clean_pv_data(inputh_path, output_path, extracted_type)
result_df = mock_save_df.call_args[0][0]
assert "January" in result_df.columns
assert "February" in result_df.columns
assert result_df.loc["2022", "January"] == 100
assert result_df.loc["2022", "February"] == 150
assert "Net nominal capacity" not in result_df.values
@@ -0,0 +1,21 @@
""""
Unit test verifying that the clip function correctly clips spatial data
using a provided overlay.
"""
import geopandas as gpd
from clip import clip
def test_clip(local_tmp_path): # pylint: disable=redefined-outer-name
"""Test that clip() writes a clipped GeoDataFrame matching the expected result"""
input_path = "tests/data/solarparcs.gpkg"
overlay_path = "tests/data/germany.gpkg"
expected_path = "tests/data/clip.gpkg"
output_path = local_tmp_path / "output.gpkg"
clip(input_path, overlay_path, output_path)
expected_data = gpd.read_file(expected_path)
output_data = gpd.read_file(output_path)
assert len(expected_data) == len(output_data)
@@ -0,0 +1,53 @@
"""
Unit tests for the collapse_columns function, which merges multiple columns
by summing their values.
"""
from unittest.mock import patch
import pandas as pd
from collapse_columns import collapse_columns
@patch("utils.save_df")
@patch("utils.read_df")
@patch("checks.check_dir")
@patch("checks.check_path")
def test_sum_values_correctly(_mock_check_path, _mock_check_dir,
mock_read_df, mock_save_df):
"""
Test that collapse_columns sums values across all columns and stores the
result in the specified column.
"""
input_df = pd.DataFrame({"A": [1, 2, 3], "B": [4, 5, 6]})
column_name = "sum"
expected_df = pd.DataFrame({column_name: [5, 7, 9]})
mock_read_df.return_value = input_df
collapse_columns("in.csv", "out.csv", column_name)
result_df = mock_save_df.call_args[0][0]
pd.testing.assert_frame_equal(result_df, expected_df)
@patch("utils.save_df")
@patch("utils.read_df")
@patch("checks.check_dir")
@patch("checks.check_path")
def test_output_dimension(_mock_check_path,
_mock_check_dir, mock_read_df, mock_save_df):
"""
Test that the output DataFrame has one column after merging and retains
the original number of rows.
"""
input_df = pd.DataFrame({
"X": [1, 2],
"Y": [3, 4]
})
column_name = "sum"
mock_read_df.return_value = input_df
expected_rows = input_df.shape[0]
collapse_columns("in.csv", "out.csv", column_name)
result_df = mock_save_df.call_args[0][0]
assert result_df.shape[1] == 1, "Output should have exactly one column"
assert result_df.shape[0] == expected_rows, "Output should have same number of rows as input"
@@ -0,0 +1,66 @@
"""
Unit tests for logging functions.
"""
from unittest.mock import patch
import logs
def test_log_processing():
"""Test logging of processing start."""
with patch('logging.info') as mock_info:
logs.log_processing("ToolA")
mock_info.assert_called_once_with("Processing %s:", "ToolA")
def test_log_processed():
"""Test logging of processing completion."""
with patch('logging.info') as mock_info:
logs.log_processed("ToolB")
mock_info.assert_called_once_with("Processed: %s", "ToolB")
def test_log_read():
"""Test logging of file read event."""
with patch('logging.info') as mock_info:
logs.log_read("/path/to/file.txt")
mock_info.assert_called_once_with("Read: %s", "/path/to/file.txt")
def test_log_read_error():
"""Test logging of file read error."""
with patch('logging.error') as mock_error:
logs.log_read_error("/path/to/missing_file.txt")
mock_error.assert_called_once_with(
"Failed to read: %s", "/path/to/missing_file.txt")
def test_log_saved():
"""Test logging of file save event."""
with patch('logging.info') as mock_info:
logs.log_saved("/path/to/save_location.txt")
mock_info.assert_called_once_with(
"Saved: %s", "/path/to/save_location.txt")
def test_log_saved_error():
"""Test logging of file save error."""
with patch('logging.error') as mock_error:
logs.log_saved_error("/path/to/save_location.txt")
mock_error.assert_called_once_with(
"Failed to save: %s", "/path/to/save_location.txt")
def test_log_parameter():
"""Test logging of parameter."""
with patch('logging.info') as mock_info:
logs.log_parameter('parameter_name', str(999))
mock_info.assert_called_once_with(
"Tool runs with %s: %s", "parameter_name", str(999)
)
def test_log_intensive():
"""Test logging warning for ressource intensiveness."""
with patch('logging.warning') as mock_info:
logs.log_intensive()
mock_info.assert_called_once_with("Tool is resource intensive")
@@ -0,0 +1,23 @@
"""Unit tests for the merge_series function"""
import os
import pandas as pd
from merge_series import merge_series
def test_merge_series(local_tmp_path):
"""
Test that merge_series combines the data and creates
a non empty CSV.file.
"""
input_directory = "tests/data/sunshine_duration"
output_path = local_tmp_path / "merged.csv"
expected_path = "tests/data/sunshine_duration.csv"
merge_series(input_directory, output_path)
expected_data = pd.read_csv(expected_path)
output_data = pd.read_csv(output_path)
assert os.path.exists(output_path)
assert os.path.getsize(output_path) > 0
pd.testing.assert_frame_equal(expected_data, output_data)
@@ -0,0 +1,40 @@
'''
Unit tests for the plot_change function.
These tests use mocking to isolate the function from file I/O,
logging, and plotting. They verify that the function behaves
correctly for both monthly and yearly data formats.
'''
from unittest import mock
import pandas as pd
import matplotlib
import plot_change
matplotlib.use("Agg")
@mock.patch("matplotlib.pyplot.savefig")
@mock.patch("plot_change.utils.read_df")
@mock.patch("plot_change.checks.check_empty")
@mock.patch("plot_change.checks.check_dir")
@mock.patch("plot_change.checks.check_path")
@mock.patch("plot_change.logs.log_processed")
@mock.patch("plot_change.logs.log_processing")
def test_plot_change_yearly(_mock_log_processing, _mock_log_processed,
_mock_check_path, _mock_check_dir, _mock_check_empty,
mock_read_df, mock_savefig, local_tmp_path):
"""Test yearly data case without year_filter."""
df = pd.DataFrame({
"Year": [2018, 2019, 2020],
"Electricity feed-in": [100000, 110000, 120000]
})
mock_read_df.return_value = df
input_path = "fake_input.csv"
output_path = local_tmp_path / "output_yearly.png"
title = "Yearly Plot"
plot_change.plot_change(str(input_path), str(
output_path), title, year_filter=None)
mock_read_df.assert_called_once_with(input_path, ",", 0, None)
mock_savefig.assert_called_once_with(str(output_path))
@@ -0,0 +1,17 @@
"""Unit tests for the plot_geo module."""
import os
from plot_geo import plot_geo
def test_plot_geo(local_tmp_path): # pylint: disable=redefined-outer-name
"""Test that gpkg2png generates a non-empty PNG file"""
data_path = "tests/data/solarparcs.gpkg"
base_path = "tests/data/germany.gpkg"
output_path = local_tmp_path / "output.png"
title = "Test Map"
plot_geo(data_path, base_path, output_path, title)
assert os.path.exists(output_path)
assert os.path.getsize(output_path) > 0
@@ -0,0 +1,19 @@
"""Unit tests for the polygons2area function in src.polygons2area.py"""
import os
from polygons2area import polygons2area
def test_polygons2area(local_tmp_path): # pylint: disable=redefined-outer-name
"""
Test that polygons2area computes the area of input polygons and writes the result
to a non-empty output text file.
"""
input_path = "tests/data/clip.gpkg"
output_path = local_tmp_path / "output.txt"
unit = "m"
polygons2area(input_path, output_path, unit)
assert os.path.exists(output_path)
assert os.path.getsize(output_path) > 0
@@ -0,0 +1,25 @@
"""
Unit tests for the trim module's trimming functionality.
"""
import pandas as pd
import trim
def test_trim_function(local_tmp_path):
"""Test trimming rows from a CSV file using trim.trim."""
input_csv = local_tmp_path / "input.csv"
output_csv = local_tmp_path / "output.csv"
sample_data = pd.DataFrame({
"id": [1, 2, 3, 4, 5],
"name": ["Alice", "Bob", "Charlie", "David", "Eve"],
"value": [10, 20, 30, 40, 50]
})
sample_data.to_csv(input_csv, index=False)
trim.trim(str(input_csv), str(output_csv), "1", "2")
result = pd.read_csv(output_csv)
expected = sample_data.iloc[1:-2]
pd.testing.assert_frame_equal(result.reset_index(
drop=True), expected.reset_index(drop=True))
@@ -0,0 +1,94 @@
"""Unit tests for utils.py"""
from unittest.mock import Mock
import pytest
import pandas as pd
import geopandas as gpd
import utils
def test_is_number():
"""Test if strings are correctly identified as numbers or not."""
assert utils.is_number("42") is True
assert utils.is_number("3.14") is True
assert utils.is_number("abc") is False
assert utils.is_number("") is False
def test_parse_number():
"""Test parsing numeric strings to int or float and raising errors for invalid input."""
assert utils.parse_number("42") == 42
assert utils.parse_number("3.14") == 3.14
with pytest.raises(ValueError):
utils.parse_number("not a number")
def test_request_response():
"""Test HTTP GET request and check if response has expected structure."""
response = utils.request_response("https://httpbin.org/get")
assert response.status_code == 200
assert "url" in response.json()
def test_parse_html():
"""Test parsing HTML content and extracting elements."""
mock_response = Mock()
mock_response.text = "<html><body><p>Hello</p></body></html>"
result = utils.parse_html(mock_response)
assert result.p.text == "Hello"
def test_read_file(local_tmp_path):
"""Test reading content from a file."""
test_path = local_tmp_path / "test.txt"
test_path.write_text("test content")
content = utils.read_file(test_path)
assert content == "test content"
def test_save_file(local_tmp_path):
"""Test saving content to a file."""
test_path = local_tmp_path / "output.txt"
utils.save_file("saved content", test_path)
assert test_path.read_text() == "saved content"
def test_read_df(local_tmp_path):
"""Test reading a CSV file into a pandas DataFrame."""
csv_path = local_tmp_path / "test.csv"
csv_path.write_text("a,b\n1,2\n3,4")
df = utils.read_df(csv_path, seperator=",", header_line=0, icol=None)
assert not df.empty
def test_save_df(local_tmp_path):
"""Test saving a pandas DataFrame to a CSV file."""
df = pd.DataFrame({"x": [1, 2], "y": [3, 4]})
path = local_tmp_path / "output.csv"
utils.save_df(df, path)
loaded = pd.read_csv(path, index_col=0)
pd.testing.assert_frame_equal(df, loaded)
def test_read_gdf(local_tmp_path):
"""Test reading a GeoDataFrame from a GeoPackage file."""
gdf = gpd.GeoDataFrame(
{'col': [1]}, geometry=gpd.points_from_xy([0], [0]), crs="EPSG:4326")
path = local_tmp_path / "test.gpkg"
gdf.to_file(path, driver="GPKG")
result = utils.read_gdf(path)
assert isinstance(result, gpd.GeoDataFrame)
assert result.equals(gdf)
def test_save_gdf(local_tmp_path):
"""Test saving a GeoDataFrame to a GeoPackage file."""
gdf = gpd.GeoDataFrame(
{'col': [1]}, geometry=gpd.points_from_xy([0], [0]), crs="EPSG:4326")
path = local_tmp_path / "output.gpkg"
utils.save_gdf(gdf, path)
loaded = gpd.read_file(path)
assert isinstance(loaded, gpd.GeoDataFrame)
assert loaded.equals(gdf)
@@ -0,0 +1,97 @@
"""Unit test suite for the `wgetdir` module."""
from urllib.parse import urljoin
from unittest import mock
from bs4 import BeautifulSoup
import wgetdir
def test_parse_urls():
"""Test that `parse_urls` correctly extracts valid file URLs from various HTML inputs."""
html = '''
<html>
<body>
<a href="https://example.com/file1.txt">File 1</a>
<a href="https://example.com/folder/">Folder</a>
<a href="/local/path/">Local Path</a>
<a href="../">Parent</a>
<a href="/">Root</a>
<a href="https://example.com/file2.pdf">File 2</a>
</body>
</html>
'''
soup = BeautifulSoup(html, 'html.parser')
soup_empty = BeautifulSoup("", 'html.parser')
html_missing = '''
<html><body>
<a>Missing href</a>
<a href="">Empty href</a>
</body></html>
'''
soup_missing = BeautifulSoup(html_missing, 'html.parser')
html_wrong_tag = '''
<html><body>
<div href="https://shouldnot.be.included">Wrong tag</div>
</body></html>
'''
soup_wrong_tag = BeautifulSoup(html_wrong_tag, 'html.parser')
result1 = wgetdir.parse_urls(soup)
result2 = wgetdir.parse_urls(soup_empty)
result3 = wgetdir.parse_urls(soup_missing)
result4 = wgetdir.parse_urls(soup_wrong_tag)
assert result1 == [
"https://example.com/file1.txt",
"https://example.com/file2.pdf"
]
assert not result2
assert result3 == [""]
assert not result4
def test_wget(local_tmp_path):
"""Test that `wget` downloads a file from a URL."""
url = "https://example.com/data.txt"
expected_file = local_tmp_path / "data.txt"
fake_content = "This is fake downloaded content."
with mock.patch("wgetdir.checks.check_empty") as _, \
mock.patch("wgetdir.checks.check_dir") as _, \
mock.patch("wgetdir.utils.request_response") as mock_request:
mock_response = mock.Mock()
mock_response.text = fake_content
mock_request.return_value = mock_response
wgetdir.wget(url, str(local_tmp_path))
assert expected_file.exists(), "File was not created"
with expected_file.open("r") as f:
content = f.read()
assert content == fake_content, "File content does not match"
def test_wgetdir(local_tmp_path):
"""Test that `wgetdir` processes URL by retrieving all contained files."""
directory_url = "https://example.com/data/"
urls = ["file1.txt", "file2.pdf"]
full_urls = [urljoin(directory_url, u) for u in urls]
with mock.patch("wgetdir.logs.log_processing"), \
mock.patch("wgetdir.logs.log_processed"), \
mock.patch("wgetdir.checks.check_empty"), \
mock.patch("wgetdir.checks.check_dir"), \
mock.patch("wgetdir.utils.request_response") as mock_response, \
mock.patch("wgetdir.utils.parse_html") as _, \
mock.patch("wgetdir.parse_urls") as mock_parse_urls, \
mock.patch("wgetdir.wget") as mock_wget:
mock_parse_urls.return_value = urls
mock_response.return_value = mock.Mock()
wgetdir.wgetdir(directory_url, str(local_tmp_path))
calls = [mock.call(url, str(local_tmp_path)) for url in full_urls]
mock_wget.assert_has_calls(calls, any_order=False)
@@ -0,0 +1,102 @@
# ----------------------------------------------------- #
# Calculate Energy and Analysis Workflow #
# ----------------------------------------------------- #
# Processes photovoltaic and solar datasets to calculate and calculate yearly energy
rule calculate_theoretical_energy:
input:
area = f"{RESULTS}/solarparcs_area.txt",
sunshine_duration = f"{RESULTS}/sunshine_duration.csv"
output:
f"{RESULTS}/solarparcs_energy.csv"
shell:
f"{PYTHON} {SRC}/calculate_energy.py {{input.area}} {{input.sunshine_duration}} {{output}} {EFFICIENCY}"
# Sums each row across all columns of the solaparc data and writes the result as a single-column file
rule collapse_columns_theoretical_energy:
input:
f"{RESULTS}/solarparcs_energy.csv"
output:
f"{RESULTS}/solarparcs_energy_yearly.csv"
params:
column_name = "Energy [kWh]"
shell:
f'{PYTHON} {SRC}/collapse_columns.py {{input}} {{output}} "{{params.column_name}}"'
# trim rows from the beginning and end of the solarparc CSV file and save the result
rule trim_theoretical_energy:
input:
f"{RESULTS}/solarparcs_energy_yearly.csv"
output:
f"{RESULTS}/solarparcs_energy_yearly_trimmed.csv"
params:
arg1 = "67",
arg2 = "0"
shell:
f"{PYTHON} {SRC}/trim.py {{input}} {{output}} {{params.arg1}} {{params.arg2}}"
# plots the yearly changes of solarparc energy
rule plot_yearly_energy_change:
input:
f"{RESULTS}/solarparcs_energy_yearly_trimmed.csv"
output:
f"{REPORT}/solarparcs_energy_yearly_trimmed.png"
params:
title = "Theoretical Energy 2018-2024",
year_filter=None
shell:
f'{PYTHON} {SRC}/plot_change.py {{input}} {{output}} "{{params.title}}" {{params.year_filter}}'
# plots the yearly changes of phtovoltaic energy
rule plot_yearly_change_photovoltaic:
input:
f"{RESULTS}/pv_data_cleaned_trimmed_collapsed.csv"
output:
f"{REPORT}/pv_data_cleaned_trimmed_collapsed.png"
params:
title = "Photovoltaic",
year_filter=None
shell:
f'{PYTHON} {SRC}/plot_change.py {{input}} {{output}} "{{params.title}}" {{params.year_filter}}'
# Calculates and compares yearly energy sums from actual and theoretical solar data
rule calculate_difference:
input:
actual = f"{RESULTS}/pv_data_cleaned_trimmed_collapsed.csv",
theoretical = f"{RESULTS}/solarparcs_energy_yearly_trimmed.csv"
output:
f"{RESULTS}/energy_difference.csv"
shell:
f"{PYTHON} {SRC}/calculate_difference.py {{input.actual}} {{input.theoretical}} {{output}}"
# plots the monthly change of photovoltaic energy
rule plot_monthly_change_photovoltaic:
input:
f"{RESULTS}/pv_data_cleaned.csv"
output:
f"{REPORT}/pv_monthly_change_{YEAR}.png"
params:
title = "'Monthly Change Photovoltaic'",
year_filter={YEAR}
shell:
f'{PYTHON} {SRC}/plot_change.py {{input}} {{output}} "{{params.title}}" "{{params.year_filter}}"'
# Plots yearly differences between actual and theoretical photovoltaic energy
rule plot_energy difference:
input:
f"{RESULTS}/energy_difference.csv"
output:
f"{REPORT}/energy_difference.png"
params:
title = "Yearly Energy Difference: Actual – Theoretical (kWh)",
year_filter=None
shell:
f'{PYTHON} {SRC}/plot_change.py {{input}} {{output}} "{{params.title}}" {{params.year_filter}}'
@@ -0,0 +1,41 @@
# ----------------------------------------------------- #
# Processing of Photovoltaic Data Workflow #
# ----------------------------------------------------- #
# Processes and cleans the photovoltaic CSV dataset, then extracts specific data types "Electricity feed-in systems"
rule load_and_clean_photovoltaic_data:
input:
csv_in = f"{DATA}/pv_data.csv"
output:
csv_out = f"{RESULTS}/pv_data_cleaned.csv"
params:
label = "Electricity feed-in"
shell:
r'"{PYTHON}" "{SRC}/clean_pv_data.py" "{input.csv_in}" "{output.csv_out}" "{params.label}"'
# trim rows from the beginning and end of the photovoltaic CSV file and save the result
rule trim_photovoltaic:
input:
f"{RESULTS}/pv_data_cleaned.csv"
output:
f"{RESULTS}/pv_data_cleaned_trimmed.csv"
params:
arg1 = "0",
arg2 = "1"
shell:
f"{PYTHON} {SRC}/trim.py {{input}} {{output}} {{params.arg1}} {{params.arg2}}"
# Sums each row across all columns of the photovoltaic data and writes the result as a single-column file
rule collapse_columns_photovoltaic:
input:
csv_in = f"{RESULTS}/pv_data_cleaned_trimmed.csv"
output:
csv_out = f"{RESULTS}/pv_data_cleaned_trimmed_collapsed.csv"
params:
column_name = "Energy [kWh]"
shell:
f'{PYTHON} "{SRC}/collapse_columns.py" "{{input.csv_in}}" "{{output.csv_out}}" "{{params.column_name}}"'
@@ -0,0 +1,35 @@
# ----------------------------------------------------- #
# Processing of Solarparc Data Workflow #
# ----------------------------------------------------- #
# performs spatial clipping of Geopackages using GeoPandas
rule load_solarparc_and_border_data:
input:
solarparcs = f"{DATA}/solarparcs.gpkg",
germany = f"{DATA}/germany.gpkg"
output:
f"{RESULTS}/solarparcs_clipped.gpkg"
shell:
f"{PYTHON} {SRC}/clip.py {{input.solarparcs}} {{input.germany}} {{output}}"
# Reads geographic data, plots it on a base map, and saves the visualization as an image
rule plot_solarparc_map:
input:
clipped = f"{RESULTS}/solarparcs_clipped.gpkg",
germany = f"{DATA}/germany.gpkg"
output:
f"{REPORT}/solarparcs.png"
shell:
f'{PYTHON} {SRC}/plot_geo.py {{input.clipped}} {{input.germany}} {{output}} "Solarparcs in Germany 2025"'
# Calculate and compare yearly energy sums from actual and theoretical solar data
rule calculate_area:
input:
f"{RESULTS}/solarparcs_clipped.gpkg"
output:
f"{RESULTS}/solarparcs_area.txt"
shell:
f"{PYTHON} {SRC}/polygons2area.py {{input}} {{output}} m"
@@ -0,0 +1,59 @@
# ----------------------------------------------------- #
# Processing of Sunshine Duration Data Workflow #
# ----------------------------------------------------- #
# Downloads all files from a URL to a local directory
rule load_sunshine_duration_data:
output:
directory(f"{RESULTS}/{SUNSHINE_DATA}/")
shell:
f'{PYTHON} {SRC}/wgetdir.py "https://opendata.dwd.de/climate_environment/CDC/regional_averages_DE/monthly/sunshine_duration/" {{output}}'
# Combines twelve CSV files from a directory by selecting German averages into one dataframe
rule merge_series:
input:
f"{RESULTS}/{SUNSHINE_DATA}/"
output:
f"{RESULTS}/sunshine_duration.csv"
shell:
f"{PYTHON} {SRC}/merge_series.py {{input}} {{output}}"
# Sums each row across all columns of the sunshiine duration data and write the result as a single-column file
rule collapse_columns_sunshine_duration:
input:
f"{RESULTS}/sunshine_duration.csv"
output:
f"{RESULTS}/sunshine_duration_yearly.csv"
params:
column_name = "Duration [h]"
shell:
f'{PYTHON} {SRC}/collapse_columns.py {{input}} {{output}} "{{params.column_name}}"'
# trim rows from the beginning and end of the solarparc CSV file and save the result
rule trim_sunshine_duration:
input:
f"{RESULTS}/sunshine_duration_yearly.csv"
output:
f"{RESULTS}/sunshine_duration_yearly_trimmed.csv"
params:
arg1 = "67",
arg2 = "0"
shell:
f"{PYTHON} {SRC}/trim.py {{input}} {{output}} {{params.arg1}} {{params.arg2}}"
# plots the yearly changes of sunshine duration
rule plot_yearly_change_sunshine:
input:
f"{RESULTS}/sunshine_duration_yearly_trimmed.csv"
output:
f"{REPORT}/sunshine_duration_yearly_trimmed.png"
params:
title = "Sunshine Duration 2018-2024",
year_filter = None
shell:
f'{PYTHON} {SRC}/plot_change.py {{input}} {{output}} "{{params.title}}" {{params.year_filter}}'
@@ -0,0 +1,67 @@
"""
Calculate and compare yearly energy sums from actual photovoltaic and theoretical solar datasets.
This script reads two CSV datasets containing monthly energy values per year — one representing
actual photovoltaic energy measurements, the other representing theoretical solar radiation
estimates. It sums the monthly values per year, and merges the results into a single output file
for comparison or further analysis.
Usage:
python calculate_difference.py <actual_energy_path> <theoretical_energy_path>
<output_path>
Arguments:
actual_energy_path (str or Path): path to CSV file with actual photovoltaic energy data.
theoretical_energy_path (str or Path): path to CSV file with theoretical solar radiation
data.
output_path (str or Path): destination path for the merged yearly sums output
CSV.
"""
import os
import sys
import pandas as pd
import utils
import checks
import logs
def calc_energy_difference(df_actual_path, df_theoretical_path, output_path):
"""
Calculate the absolute yearly energy difference between actual and theoretical datasets.
Reads two CSV files with yearly energy data, computes the absolute difference
in energy values for each common year, and saves the result as a CSV.
Args:
df_actual_path (str or Path): Path to the CSV file with actual photovoltaic energy
data.
df_theoretical_path (str or Path): Path to the CSV file with theoretical solar radiation
data.
output_path (str or Path): Path where the output CSV with yearly energy differences
will be saved.
Returns:
None: The result is saved to `output_path`.
"""
logs.log_processing(os.path.basename(__file__))
checks.check_path(df_actual_path)
checks.check_path(df_theoretical_path)
checks.check_dir(output_path)
df_actual = utils.read_df(df_actual_path, ",", 0, None)
df_theoretical = utils.read_df(df_theoretical_path, ",", 0, None)
diff_difference = abs(
df_actual[df_actual.columns[1]] - df_theoretical[df_theoretical.columns[1]])
df_diff = pd.DataFrame({
'year': df_actual['year'],
'Energy Difference [kWh]': diff_difference})
utils.save_df(df_diff, output_path, False)
logs.log_processed(os.path.basename(__file__))
if __name__ == "__main__":
checks.check_args(4, __doc__)
calc_energy_difference(sys.argv[1], sys.argv[2], sys.argv[3])
@@ -0,0 +1,64 @@
"""
Calculate theoretical energy generation of solar parks.
This script reads the area of solar parks (in m²) and sunshine duration (in hours),
then calculates the theoretical energy produced (in kWh) using a given efficiency.
Usage:
python calculate_energy.py <area_data_path> <sunshine_data_path> <output_data_path> <efficiency>
Arguments:
area_data_path (str or Path): Path to the input file with solar park area data.
sunshine_data_path (str or Path): Path to the input file with sunshine duration data.
output_data_path (str or Path): Path where the output file will be saved.
efficiency (float): Average efficiency of PV panels (as a decimal, e.g. 0.15
for 15%)
"""
import os
import sys
import utils
import checks
import logs
def calculate_energy(area_path, sunshine_path, output_path, efficiency):
"""
Calculate the theoretical energy production of solar parks in kWh.
Reads input files with solar park area and sunshine duration, then computes
energy using the formula: energy = area * sunshine_duration * power_per_m2 * efficiency.
Args:
area_path (str or Path): Path to the input file with area data (m²).
sunshine_path (str or Path): Path to the input file with sunshine duration data (hours).
output_path (str or Path): Path to save the calculated energy output CSV.
efficiency (float): Average efficiency of PV panels (decimal).
Returns:
None: Saves the calculated energy DataFrame to the specified output path.
"""
logs.log_processing(os.path.basename(__file__))
logs.log_parameter('efficiency', str(efficiency))
checks.check_path(area_path)
checks.check_path(sunshine_path)
checks.check_dir(output_path)
checks.check_empty(efficiency)
content = utils.read_file(area_path)
area = utils.parse_number(content)
sunshine_data = utils.read_df(sunshine_path, ",", 0, 0)
efficiency_number = float(efficiency)
power_per_m2 = 0.1
# main calculation
factor = power_per_m2 * area * efficiency_number
energy = sunshine_data * factor
utils.save_df(energy, output_path)
logs.log_processed(os.path.basename(__file__))
if __name__ == "__main__":
checks.check_args(5, __doc__)
calculate_energy(sys.argv[1], sys.argv[2], sys.argv[3], sys.argv[4])
@@ -0,0 +1,176 @@
"""
Basic input validation checks.
Functions to check for :
- non-empty variables,
- existing file or directory paths
- correct command-line argument count
- two CRS's if they are equal
- dataframe valid and not empty
- if GeoDataFrame is valid and not empty
- if variable is member of a list
"""
import os
import sys
from pathlib import Path
import pandas as pd
import geopandas as gpd
def check_empty(variable):
"""
Checks if variable is not empty.
Arguments:
variable: variable to check.
Returns:
None
Raises:
ValueError: If the variable is empty.
"""
if not variable:
raise ValueError("Variable is empty")
def check_path(path):
"""
Checks if path exists.
Arguments:
path (str): path to check.
Returns:
None
Raises:
FileNotFoundError: If the path does not exist.
"""
if not os.path.exists(path):
raise FileNotFoundError(f"Invalid path: {path}")
def check_dir(path):
"""
Checks if directory exists.
Arguments:
path (str): path to check.
Returns:
None
Raises:
FileNotFoundError: If the directory's parent path does not exist.
"""
if not Path(path).parent.exists():
raise FileNotFoundError(f"Invalid directory path: {path}")
def check_args(num_arg, output):
"""
Checks sys arguments length against expected.
Args:
num_arg (int): expected number of arguments.
output (str): string to be printed in case of check failure.
Returns:
None
Raises:
RuntimeError: If the number of arguments does not match `num_arg`.
"""
if len(sys.argv) != num_arg:
raise RuntimeError(output)
def check_crs(crs1, crs2):
"""
Checks two CRS's if they are equal.
Arguments:
crs1: first CRS.
crs2: second CRS.
Returns:
None
Raises:
ValueError: If CRS's are not equal.
"""
if crs1 != crs2:
raise ValueError("CRS missmatch.")
def check_df(df):
"""
Checks if DataFrame is valid and not empty.
Arguments:
df: DataFrame to check.
Returns:
None
Raises:
TypeError: If not a DataFrame
ValueError: If DataFrame is empty
ValueError: If all values in the DataFrame are NaN
"""
if not isinstance(df, pd.DataFrame):
raise TypeError("Not a DataFrame.")
if df.empty:
raise ValueError("DataFrame is empty.")
if df.isna().all().all():
raise ValueError("DataFrame contains only NaN values.")
def check_gdf(gdf):
"""
Checks if GeoDataFrame is valid and not empty.
Arguments:
gdf: GeoDataFrame to check.
Returns:
None
Raises:
TypeError: if not an GeoDataFrame
ValueError: If GeoDataFrame is empty
ValueError: If includes empty geometries
"""
if not isinstance(gdf, gpd.GeoDataFrame):
raise TypeError("Not a GeoDataFrame.")
if gdf.empty:
raise ValueError("GeoDataFrame is empty.")
if gdf.geometry.isna().all():
raise ValueError("GeoDataFrame includes empty geometries.")
def check_member(variable, members):
"""
Check if variable is member of a list.
Arguments:
variable: to check
members (list): to check against.
Returns:
None
Raises:
ValueError: If variable is not member of allowed varibales.
"""
if variable not in members:
raise ValueError(f"Invalid value: {
variable}. Must be one of: {members}")
@@ -0,0 +1,85 @@
"""
Process and clean photovoltaic dataset from Destatis.
This script reads a CSV dataset (exported from the Destatis web portal), removes
unnecessary table structures caused by web-to-CSV conversion, and saves a cleaned version.
It also allows extracting a subset of the data, e.g. "Electricity feed-in systems",
"Net nominal capacity", or "Electricity feed-in".
Original dataset source:
https://www-genesis.destatis.de/datenbank/online/statistic/43312/table/43312-0001
Usage:
python clean_pv_data.py <dataset_path> <save_path> <extracted_type>
Arguments:
dataset_path (str or Path): Path to the raw CSV dataset file.
save_path (str or Path): Path where the cleaned CSV will be saved.
extracted_type (str): Type of data to extract (e.g., "Electricity feed-in").
"""
import os
import sys
import pandas as pd
import utils
import checks
import logs
def clean_pv_data(input_path, output_path, extracted_type):
"""
Clean and preprocess the photovoltaic dataset from Destatis.
This function selects relevant columns, renames them for clarity,
defines categorical ordering for months and data types, sorts the data,
and extracts the specified subset. The result is saved as a pivot table CSV.
Args:
input_path (str or Path): Path to the raw input CSV file.
output_path (str or Path): Path to save the cleaned and extracted CSV.
extracted_type (str): The category of data to extract (e.g., "Electricity feed-in").
Returns:
None: The processed DataFrame is saved to `output_path`.
"""
logs.log_processing(os.path.basename(__file__))
checks.check_path(input_path)
checks.check_dir(output_path)
checks.check_empty(extracted_type)
df = utils.read_df(input_path, seperator=";", header_line=0, icol=None)
df_relevant = df[["time", "1_variable_attribute_label",
"value", "value_variable_label"]]
df_cleaned = df_relevant.rename(columns={
'time': 'year',
'1_variable_attribute_label': 'month',
'value_variable_label': 'type'
})
monthly_order = ["January", "February", "March", "April", "May", "June",
"July", "August", "September", "October", "November", "December"]
# Units: [number, MW, Mwh]
type_order = ["Electricity feed-in systems",
"Net nominal capacity", "Electricity feed-in"]
df_cleaned['month'] = pd.Categorical(
df_cleaned['month'], categories=monthly_order, ordered=True)
df_cleaned['type'] = pd.Categorical(
df_cleaned['type'], categories=type_order, ordered=True)
df_cleaned = df_cleaned.sort_values(
by=['year', 'month', 'type'], ascending=True)
extracted_df = df_cleaned[df_cleaned["type"] == extracted_type]
pivot_df = extracted_df.pivot(
index="year", columns="month", values="value")
checks.check_df(pivot_df)
utils.save_df(pivot_df, output_path)
logs.log_processed(os.path.basename(__file__))
if __name__ == "__main__":
checks.check_args(4, __doc__)
clean_pv_data(sys.argv[1], sys.argv[2], sys.argv[3])
@@ -0,0 +1,59 @@
"""
Perform spatial clipping of geospatial data using GeoPandas.
This script reads an input geospatial file and an overlay file,
clips the input to the overlay boundaries, and saves the clipped
result as a GeoPackage file.
Usage:
python clip.py <input_path> <overlay_path> <output_path>
Arguments:
input_path (str or Path): Path to the input geospatial file (e.g., GeoPackage, shapefile).
overlay_path (str or Path): Path to the overlay geospatial file used for clipping.
output_path (str or Path): Path to save the clipped output as a GeoPackage (.gpkg).
"""
import os
import sys
import geopandas as gpd
import utils
import checks
import logs
def clip(input_path, overlay_path, output_path):
"""
Clips input to overlay and exports result. Includes Errorhandling.
Performs spatial clipping of input data using the geometry boundaries
from the overlay data, and exports the clipped result to a GeoPackage.
Arguments:
input_path (str or Path): Path to the input file (.gpkg).
overlay_path (str or Path): Path to the clipping layer (.gpkg).
output_path (str or Path): Path to save the clipped output (.gpkg).
Returns:
None: The processed clipped output is saved to `output_path`.
"""
logs.log_processing(os.path.basename(__file__))
logs.log_intensive()
checks.check_path(input_path)
checks.check_path(overlay_path)
checks.check_dir(output_path)
input_data = utils.read_gdf(input_path)
overlay_data = utils.read_gdf(overlay_path)
checks.check_crs(input_data.crs, overlay_data.crs)
clipped_data = gpd.clip(input_data, overlay_data)
checks.check_gdf(clipped_data)
utils.save_gdf(clipped_data, output_path)
logs.log_processed(os.path.basename(__file__))
if __name__ == "__main__":
checks.check_args(4, __doc__)
clip(sys.argv[1], sys.argv[2], sys.argv[3])
@@ -0,0 +1,46 @@
"""
This script reads a tabular CSV dataset, sums the values across all columns for each row,
and writes the resulting single-column DataFrame to an output file.
Usage:
python collapse_columns.py <input_path> <output_path> <column_name>
Arguments:
input_path (str or Path): Path to the input CSV file containing the original tabular data.
output_path (str or Path): Path where the collapsed output CSV file will be saved.
column_name (str): Name of the column in the output file representing the summed values.
"""
import os
import sys
import utils
import checks
import logs
def collapse_columns(input_path, output_path, column_name):
"""
Reads a tabular dataset and collapses each row into a single value by summing across columns.
The result is written as a single-column DataFrame to the output path.
Args:
input_path (str or Path): Path to the input CSV file.
output_path (str or Path): Path where the output CSV file will be saved.
column_name (str): Name of the resulting summed column.
Returns:
None: The resulting summed dataframe is saved to `output_path`
"""
logs.log_processing(os.path.basename(__file__))
logs.log_parameter("column_name", str(column_name))
checks.check_path(input_path)
checks.check_dir(output_path)
data = utils.read_df(input_path, ",", 0, 0)
data = data.sum(axis=1).to_frame(name=column_name)
utils.save_df(data, output_path)
logs.log_processed(os.path.basename(__file__))
if __name__ == "__main__":
checks.check_args(4, __doc__)
collapse_columns(sys.argv[1], sys.argv[2], sys.argv[3])
@@ -0,0 +1,131 @@
"""
Logging setup and helper functions for processing events.
Configures logging to file and console, with functions
to log start, completion, and saving steps of processing.
Containing logs:
- processing
- processed
- read
- read error
- saved
- saved error
- parameter
- intensive
"""
import logging
logging.basicConfig(
level=logging.INFO,
format="%(asctime)s - %(levelname)s - %(message)s",
handlers=[
logging.FileHandler("log.txt"),
logging.StreamHandler()
]
)
def log_processing(tool_name):
"""
Logs the start of processing for a given tool.
Args:
tool_name (str): Name of the tool being processed.
Returns:
None
"""
logging.info("Processing %s:", tool_name)
def log_processed(tool_name):
"""
Logs the successful completion of processing for a given tool.
Args:
tool_name (str): Name of the tool that has been processed.
Returns:
None
"""
logging.info("Processed: %s", tool_name)
def log_read(path):
"""
Logs the reading of a file from a given path.
Arguments:
path (str): Path to file which were read.
Returns:
None
"""
logging.info("Read: %s", path)
def log_read_error(path):
"""
Logs the reading error of a given path.
Arguments:
path (str): Path were the file should be read.
Returns:
None
"""
logging.error("Failed to read: %s", path)
def log_saved(path):
"""
Logs the saving of output to a given path.
Args:
path (str): Path where the output was saved.
Returns:
None
"""
logging.info("Saved: %s", path)
def log_saved_error(path):
"""
Logs the saving error of output to a given path.
Arguments:
path (str): Path were the output should be saved.
Returns:
None
"""
logging.error("Failed to save: %s", path)
def log_parameter(parameter_name, parameter_value):
"""
Logs a parameter.
Arguments:
parameter_name (str): Parameter name with which the tool is called.
parameter_value (str): Parameter value with which the tool is called.
Returns:
None
"""
logging.info("Tool runs with %s: %s", parameter_name, parameter_value)
def log_intensive():
"""
Logs warning for resource intensiveness.
Arguments:
None
Returns:
None
"""
logging.warning("Tool is resource intensive")
@@ -0,0 +1,65 @@
"""
This script combines monthly sunshine duration data from twelve CSV files
in a given directory by extracting the German average values and merging
them into a single DataFrame. The combined data is then saved as a CSV file.
Original datasets from:
https://opendata.dwd.de/climate_environment/CDC/regional_averages_DE/monthly/sunshine_duration/
Usage:
python merge_series.py <input_directory> <output_path>
Arguments:
input_directory (Path or str): Directory containing the twelve monthly CSV input files.
output_path (Path or str): Path where the combined CSV output file will be saved.
"""
import os
import sys
import pandas as pd
import utils
import checks
import logs
def merge_series(input_directory, output_path):
"""
Extracts the German average sunshine duration series from twelve monthly CSV files
in the specified directory and combines them into a single pandas DataFrame.
The DataFrame columns represent months, and rows represent years.
Args:
input_directory (Path or str): Directory containing the input CSV files.
save_path (Path or str): File path to save the combined CSV output.
Returns:
None: The resulting combined dataframe is saved to `output_path`
"""
logs.log_processing(os.path.basename(__file__))
checks.check_path(input_directory)
checks.check_dir(output_path)
files = os.listdir(input_directory)
checks.check_empty(files)
txt_files = sorted([f for f in files if f.lower().endswith('.txt')])
series_list = []
month_list = ["January", "February", "March", "April", "May", "June",
"July", "August", "September", "October", "November", "December"]
for file, month in zip(txt_files, month_list):
data = utils.read_df(input_directory+"/"+file, ";", 1, 0)
data = data["Deutschland"]
data.name = month
series_list.append(data)
df = pd.concat(series_list, axis=1)
df = df[df.index != 2025]
utils.save_df(df, output_path)
logs.log_processed(os.path.basename(__file__))
if __name__ == "__main__":
checks.check_args(3, __doc__)
merge_series(sys.argv[1], sys.argv[2])
@@ -0,0 +1,83 @@
'''
This script visualizes changes in energy production or sunshine duration,
either monthly or yearly, using line plots based on a CSV input file.
Depending on the structure of the input data, the script creates:
- A monthly line plot for a selected year if the CSV contains one row per year
and one column per month.
- A yearly line plot if the CSV contains only two columns: year and value.
Usage:
python plot_change.py <input_path> <output_path> <title> <year_filter>
Arguments:
input_path (Path or str): Path to the input CSV file containing energy or sunshine data.
output_path (Path or str): File path where the generated plot (e.g., PNG) will be saved.
title (str): Title of the plot.
year_filter (str or int): Year to select from the dataset for monthly plots.
Required if the input contains monthly values.
'''
import os
import sys
import matplotlib.pyplot as plt
import utils
import checks
import logs
def plot_change(input_path, output_path, title, year_filter=None):
"""
Generate and save a line plot showing energy or sunshine changes over time.
If the input contains multiple columns (i.e., monthly values), this function will
extract the row corresponding to `year_filter` and plot the monthly trend.
Otherwise, it assumes the data is already in a year-to-value format.
Args:
input_path (Path or str): Path to the input CSV file.
output_path (Path or str): Path to save the generated plot image.
title (str): Title of the plot.
year_filter (str or int, optional): Year to filter for when plotting monthly data.
Returns:
None: The generated diagram is saved as an image file at output_path.
"""
logs.log_processing(os.path.basename(__file__))
checks.check_path(input_path)
checks.check_dir(output_path)
checks.check_empty(title)
data = utils.read_df(input_path, ",", 0, None)
if data.shape[1] > 2:
x_label = "Month"
y_label = "Energy [kWh]"
data = data.set_index("year")
x_ticks = data.columns
x_values = range(len(data.columns))
y_value = data.transpose()[int(year_filter)]
else:
x_label = data.columns[0]
y_label = data.columns[-1]
x_ticks = data.iloc[:, 0].values.astype(str)
x_values = range(len(x_ticks))
y_value = data.iloc[:, 1].values
figure, ax = plt.subplots()
ax.set_title(title)
ax.set_xlabel(x_label)
ax.set_ylabel(y_label)
ax.set_xticks(range(len(x_ticks)))
ax.set_xticklabels(x_ticks, rotation=45)
ax.plot(x_values, y_value, 'o:')
figure.tight_layout()
plt.savefig(output_path)
logs.log_processed(os.path.basename(__file__))
if __name__ == "__main__":
checks.check_args(5, __doc__)
plot_change(sys.argv[1], sys.argv[2], sys.argv[3], sys.argv[4])
@@ -0,0 +1,61 @@
"""
This script reads geographic data files, creates a map visualization by
plotting the data on top of a base map, and saves the resulting map as
an image file.
Usage:
python plot_geo.py <input_path> <base_path> <output_path> <title>
Arguments:
input_path (str or Path): Path to the geographic data file to be plotted.
base_path (str or Path): Path to the base map geographic data file.
output_path (str or Path): Path where the generated map image will be saved.
title (str): Title for the map.
"""
import os
import sys
import matplotlib.pyplot as plt
import utils
import checks
import logs
def plot_geo(input_path, base_path, output_path, title):
"""
Generate a map image by plotting geographic data over a base map.
Arguments:
input_path (str): Path to the geographic data file to be plotted.
base_path (str): Path to the base map geographic data file.
output_path (str): Path where the generated map image will be saved.
title (str): Title of the map for the plot.
Returns:
None: The generated map is saved as an image file at output_path.
"""
logs.log_processing(os.path.basename(__file__))
logs.log_parameter("title", str(title))
checks.check_path(input_path)
checks.check_path(base_path)
checks.check_dir(output_path)
checks.check_empty(title)
data = utils.read_gdf(input_path)
base = utils.read_gdf(base_path)
_, ax = plt.subplots()
ax.set_title(title)
ax.set_xlabel("Longitude")
ax.set_ylabel("Latitude")
base.plot(ax=ax, color="white", edgecolor="black")
data.plot(ax=ax, color="blue", edgecolor="blue")
plt.savefig(output_path)
logs.log_saved(output_path)
logs.log_processed(os.path.basename(__file__))
if __name__ == "__main__":
checks.check_args(5, __doc__)
plot_geo(sys.argv[1], sys.argv[2], sys.argv[3], sys.argv[4])
@@ -0,0 +1,66 @@
"""
This script calculates the total area of all geometries in a geospatial file.
It uses the appropriate UTM projection to ensure accurate area calculation.
The result is written to a plain text file in either square kilometers (default)
or square meters, depending on the specified unit.
Usage:
python polygons2area.py <input_path> <output_path> <unit>
Arguments:
input_path (str or Path): Path to the input geospatial file (e.g., GeoPackage, Shapefile).
output_path (str or Path): Path to the file where the total area will be saved as a string.
unit (str): Unit for output area: 'km' (square kilometers) or
'm' (square meters).
"""
import os
import sys
import utils
import checks
import logs
def polygons2area(input_path, output_path, unit):
"""
Computes the total area of polygons in a geospatial file and saves the result.
The geometry is first projected into the appropriate UTM CRS for accurate area calculation.
The total area is then summed and converted to the specified unit.
Args:
input_path (str or Path): Path to the input geospatial file.
output_path (str or Path): Path to the output text file.
unit (str): Unit for the area value: 'km' for square kilometers or 'm' for square meters.
Returns:
None: The total area is written as a string to the specified output_path.
"""
logs.log_processing(os.path.basename(__file__))
logs.log_parameter("unit", str(unit))
checks.check_path(input_path)
checks.check_dir(output_path)
checks.check_empty(unit)
checks.check_member(unit, ["m", "km"])
data = utils.read_gdf(input_path)
utm_crs = data.estimate_utm_crs()
projected = data.to_crs(utm_crs)
area_m2 = projected.area.sum()
if unit == "km":
area = area_m2 / 1_000_000
else:
area = area_m2
area_str = f"{area:.2f}"
utils.save_file(area_str, output_path)
logs.log_processed(os.path.basename(__file__))
if __name__ == "__main__":
checks.check_args(4, __doc__)
polygons2area(sys.argv[1], sys.argv[2], sys.argv[3])
@@ -0,0 +1,58 @@
"""
This script trims rows from the beginning and end of a CSV file and saves the result.
It reads a CSV file into a pandas DataFrame, removes a specified number of rows from
the top and bottom, and writes the trimmed DataFrame to a new file.
Usage:
python trim.py <input_path> <output_path> <beginning> <ending>
Arguments:
input_path (str or Path): Path to the input CSV file.
output_path (str or Path): Path where the trimmed CSV will be saved.
beginning (int or str): Number of rows to trim from the beginning (must be >= 0).
ending (int or str): Number of rows to trim from the end (must be >= 0).
"""
import os
import sys
import utils
import checks
import logs
def trim(input_path, output_path, beginning, ending):
"""
Trim rows from the beginning and end of a DataFrame loaded from a CSV file.
The function reads the data, trims the specified number of rows from both
the top and bottom, and saves the result to a CSV file.
Args:
input_path (str): Path to the input CSV file.
output_path (str): Path where the trimmed CSV file will be saved.
beginning (int or str): Number of rows to remove from the start.
ending (int or str): Number of rows to remove from the end.
Returns:
None: The trimmed data is written to the file specified by `output_path`.
"""
logs.log_processing(os.path.basename(__file__))
logs.log_parameter("beginning", str(beginning))
logs.log_parameter("ending", str(ending))
checks.check_path(input_path)
checks.check_dir(output_path)
checks.check_empty(beginning)
checks.check_empty(ending)
beginning = int(beginning)
ending = int(ending)
data = utils.read_df(input_path, ",", 0, 0)
data = data.iloc[beginning:len(data)-ending]
utils.save_df(data, output_path)
logs.log_processed(os.path.basename(__file__))
if __name__ == "__main__":
checks.check_args(5, __doc__)
trim(sys.argv[1], sys.argv[2], sys.argv[3], sys.argv[4])
@@ -0,0 +1,216 @@
"""
Utility functions for solalytics.
"""
import requests
import pandas as pd
import geopandas as gpd
from bs4 import BeautifulSoup
import checks
import logs
def is_number(s):
"""
Checks if given string is a number.
Arguemtns:
s (str): String to be checked.
Returns:
Bool: if string is number or not.
"""
try:
float(s)
return True
except ValueError:
return False
def parse_number(content):
"""
Parses number from file content.
Arguments:
content (str): file content to be parsed.
Returns:
float/int: parsed number.
"""
content = content.strip()
if is_number(content):
return float(content) if '.' in content else int(content)
raise ValueError("The file does not contain a valid number.")
def request_response(url):
"""
Request a reponse from a URL.
Arguments:
url (str): URL to be requested.
Returns:
response (str): response from the URL.
"""
checks.check_empty(url)
response = requests.get(url, timeout=2.5)
checks.check_empty(response)
return response
def parse_html(response):
"""
Parses HTML from a web response.
Arguments:
response (str): Response to be parsed.
Returns:
html: parsed HTML.
"""
html = BeautifulSoup(response.text, 'html.parser')
checks.check_empty(html)
return html
def read_file(path):
"""
Read a file and return its contents.
Args:
path (str): Path to the input file.
Returns:
str: The loaded content.
Raises:
FileNotFoundError: If the file does not exist.
OSError: If the file cannot be opened or read.
"""
checks.check_path(path)
try:
with open(path, 'r', encoding="utf-8") as file:
content = file.read().strip()
checks.check_empty(content)
logs.log_read(path)
return content
except OSError:
logs.log_read_error(path)
raise
def save_file(file, path):
"""
Saves fiven file as file to the specified path.
Args:
file: content of the file
path: The target file path where the file will be written.
"""
checks.check_empty(file)
checks.check_dir(path)
try:
with open(path, 'w', encoding="utf-8") as f:
f.write(file)
logs.log_saved(path)
except OSError:
logs.log_saved_error(path)
raise
def read_df(path, seperator, header_line, icol):
"""
Read a csv file and return a DataFrame.
Args:
path (str): Path to the input file.
seperator (str): csv seperator
Returns:
dataframe: The loaded content.
Raises:
FileNotFoundError: If the file does not exist.
ValueError: If dataframe is empty
OSError: If the file cannot be opened or read.
"""
checks.check_path(path)
try:
data = pd.read_csv(path, sep=seperator,
header=header_line, index_col=icol)
checks.check_df(data)
logs.log_read(path)
return data
except OSError:
logs.log_read_error(path)
raise
def save_df(df, path, index_input=True):
"""
Saves the given DataFrame as a CSV file to the specified path.
Args:
df : The DataFrame to save.
path : The target file path where the CSV will be written.
Raises:
FileNotFoundError: If the target directory does not exist.
Exception: If saving the file fails.
"""
checks.check_df(df)
checks.check_dir(path)
try:
df.to_csv(path, index=index_input)
logs.log_saved(path)
except OSError:
logs.log_saved_error(path)
def read_gdf(path):
"""
Read a geopackage and return a GeoDataFrame.
Args:
path (str): Path to the input file.
Returns:
geodataframe: The loaded content.
Raises:
FileNotFoundError: If the file does not exist.
ValueError: If dataframe is empty
OSError: If the file cannot be opened or read.
"""
checks.check_path(path)
try:
data = gpd.read_file(path)
checks.check_gdf(data)
logs.log_read(path)
return data
except OSError:
logs.log_read_error(path)
raise
def save_gdf(gdf, path):
"""
Saves the given GeoDataFrame as a GeoPackage to the specified path.
Arguments:
gdf: The GeoDataFrame to save.
path: The target file path where the .gpkg will be written.
Raises:
FileNotFoundError: If the target directory does not exist.
Exception: If saving the file fails.
"""
checks.check_gdf(gdf)
checks.check_dir(path)
try:
gdf.to_file(path, driver='gpkg')
logs.log_saved(path)
except OSError:
logs.log_saved_error(path)
@@ -0,0 +1,96 @@
"""
This script downloads all files from a given online directory URL and saves them
to a specified local directory.
It parses the provided URL for downloadable files (ignores subdirectories) and writes
the content of each file into the output directory.
Usage:
python wgetdir.py <directory_url> <directory_path>
Arguments:
directory_url (str or Path): URL to an online directory.
directory_path (str or Path): Path to a local directory where files will be saved.
"""
import sys
import os
from urllib.parse import urljoin
import utils
import checks
import logs
def parse_urls(html):
'''
Fetches all valid URLs from anchor tags in an HTML document.
Arguments:
html (BeautifulSoup): The parsed HTML using BeautifulSoup.
Returns:
List[str]: A list of URLs (hrefs) found in the HTML.
'''
urls = []
for anchor in html.find_all('a', href=True):
href = anchor['href']
if href in ('../', '/'):
continue
if href.endswith('/'):
continue
urls.append(href)
return urls
def wget(url, dir_path):
"""
Download a file form a URL and saves it to a local directory.
Arguments:
url (str): The URL of the online file.
dir_path (str): The local directory path where the file will be saved.
Returns:
None: The downloaded content is written to a file in `dir_path`.
"""
checks.check_empty(url)
checks.check_dir(dir_path)
name = url.split('/')[-1]
path = os.path.join(dir_path, name)
content = utils.request_response(url)
utils.save_file(content.text, path)
def wgetdir(directory_url, directory_path):
"""
Downloads all files from an online directory URL and saves them into a local directory.
Arguments:
directory_url (str): The URL of the online directory containing files.
directory_path (str): The local directory path where files will be saved.
Returns:
None: All downloadable files in the directory are saved to `directory_path`.
"""
logs.log_processing(os.path.basename(__file__))
logs.log_parameter("directory_url", str(directory_url))
checks.check_empty(directory_url)
checks.check_dir(directory_path)
os.makedirs(directory_path, exist_ok=True)
response = utils.request_response(directory_url)
html = utils.parse_html(response)
urls = parse_urls(html)
for url in urls:
url = urljoin(directory_url, url)
wget(url, directory_path)
logs.log_processed(os.path.basename(__file__))
if __name__ == "__main__":
checks.check_args(3, __doc__)
wgetdir(sys.argv[1], sys.argv[2])
@@ -0,0 +1,63 @@
import os
import platform
from pathlib import Path
configfile: "config/config.yml"
RESULTS = config["result_path"]
SUNSHINE_DATA = "sunshine_duration"
SUNSHINE_FILES = glob_wildcards(f"{RESULTS}/{SUNSHINE_DATA}/{{file}}.txt").file
REPORT = config["report_path"]
DATA = config["data_path"]
YEAR = config["year"]
EFFICIENCY = config["efficiency"]
os.makedirs(RESULTS, exist_ok=True)
os.makedirs(f"{RESULTS}/{SUNSHINE_DATA}/", exist_ok=True)
os.makedirs(REPORT, exist_ok=True)
os.makedirs(DATA, exist_ok=True)
SRC = config["scripts_path"]
# Detect the platform and set the Python executable
PYTHON = "python3" if platform.system() != "Windows" else "python"
# include all Rules
include: Path("rules/solarparc_data_processing.smk")
include: Path("rules/photovoltaic_data_processing.smk")
include: Path("rules/sunshine_duration_data_processing.smk")
include: Path("rules/analysis.smk")
# Defines a complete workflow run
rule all:
input:
f"{RESULTS}/sunshine_duration.csv",
f"{RESULTS}/sunshine_duration_yearly.csv",
f"{RESULTS}/sunshine_duration_yearly_trimmed.csv",
f"{REPORT}/sunshine_duration_yearly_trimmed.png",
f"{RESULTS}/solarparcs_area.txt",
f"{REPORT}/solarparcs.png",
f"{RESULTS}/solarparcs_energy.csv",
f"{RESULTS}/solarparcs_energy_yearly.csv",
f"{RESULTS}/solarparcs_energy_yearly_trimmed.csv",
f"{REPORT}/solarparcs_energy_yearly_trimmed.png",
f"{RESULTS}/pv_data_cleaned.csv",
f"{RESULTS}/pv_data_cleaned_trimmed.csv",
f"{RESULTS}/pv_data_cleaned_trimmed_collapsed.csv",
f"{REPORT}/pv_data_cleaned_trimmed_collapsed.png",
f"{RESULTS}/energy_difference.csv",
f"{REPORT}/pv_monthly_change_{YEAR}.png",
f"{REPORT}/energy_difference.png"
# Removes all output files
rule clean:
run:
if platform.system() == "Windows":
os.system(f'del /Q "{RESULTS}\\{SUNSHINE_DATA}"')
os.system(f'del /Q "{RESULTS}"')
os.system(f'del /Q "{REPORT}"')
else:
os.system(f'rm -rf {RESULTS}/{SUNSHINE_DATA}')
os.system(f'rm -rf {RESULTS}')
os.system(f'rm -rf {REPORT}')