Skip to content
siam29Public

About

This project is a simplified demonstration of a real-world big data project, showcasing efficient handling of large-scale datasets using PySpark and Apache Iceberg. This repository includes 5 sample datasets to replicate real-world scenarios, focusing on memory and time optimization while writing tables in Iceberg

Resources

Stars

0 stars

Watchers

1 watching

Forks

Latest commit

 

History

3 Commits

Folders and files

Repository files navigation

notebook

About

This project is a simplified demonstration of a real-world big data project, showcasing efficient handling of large-scale datasets using PySpark and Apache Iceberg. This repository includes 5 sample datasets to replicate real-world scenarios, focusing on memory and time optimization while writing tables in Iceberg

Resources

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages