Repository navigation
Suggestions to Introduction to Data Exploration notebook #354
Description
Activity
- changed the title
[-]Suggestions for the `Data Exploration notebok`[/-][+]Suggestions to `Introduction to Data Exploration` notebook[/+]on Feb 4, 2026 In the section "Finding the limits", the plot is showing float values for Years. Make them integer. Like using:
from matplotlib.ticker import MaxNLocator plt.gca().xaxis.set_major_locator(MaxNLocator(integer=True))
In "Cleaning missing values" section, its presented that ffill() can be used for filling NaNs. We can also show present bfill() here.
In "Exercise: Complete Happiness", if one runs the solution
solution_clean_datasetwithout writing anything (only thereturn), it returns anAttributeError: 'NoneType' object has no attribute 'sort_values', without showing the solution.In "Making frames per year" section,
# time_column = 'year'is commented. Is this needed?In solution
solution_frames_with_category, the marked for South Asia it very small (similar to Unknown), is it correct or bug?In "Bonus Exercise Fixing a library" section, there is a typo: "and will through an error". Should be throw, right?
In "Bonus Exercise Fixing a library" section, I found this path: ...modifying only the file data/plotly_intro/bubbly.py until the same code... Isnt it
/data/data_exploration/bubbly.py?In "Exercise: Complete Happiness" I do not think that the 1st question is very clear. I think the sentence "Do this by initializing a DataFrame with pd.DataFrame() with a list" is confusing. Could we explain this a bit better?
Also here we are asking to merge dataframes but actually explain the merge function below, in the 'Adding regional indicator' paragraph.
In 'Exercise: Final Happiness' we ask to Merge the cleaned_happiness_df with region_df on the 'Country name' and 'year' columns. However, region_df does not have a 'year' column and in fact the proposed solution only does a merge on 'Country name'
In "Making frames per year" section,
# time_column = 'year'is commented. Is this needed?Same in 'Plotting basic scatter plot'
In "Bonus Exercise Fixing a library" section, I found this path: ...modifying only the file data/plotly_intro/bubbly.py until the same code... Isnt it
/data/data_exploration/bubbly.py?I am also a bit confused here. Which file are we pointing them too?
tutorial/my_bubbly.pyhas 637 lines, it would be quite difficult for them to read though and identify the differences.I have created this branch: https://github.com/empa-scientific-it/python-tutorial/tree/fix%2F354-data-exploration
where I have fixed some typos and added some suggestions to make the text more comprehensible.- linked a pull request that will close this issueImprove Data Exploration notebook #368
on Apr 14, 2026 The new exercise explore_dataset needs some refinement:
- function
idxmaxis used and not mentioned in the list of panda functions - output required is quite specific, make it more flexible
- function
Building frames by function: this one is confusing because only one frame is displayed, no animation yet. So maybe lets add a play button:
'layout': { 'updatemenus': [{'type': 'buttons', 'buttons': [{'label': '▶', 'method': 'animate', 'args': [None]}]}] }
just to see the difference
Also there is a get_scatter_figure... used every so often, there should be a comment, that this is boilerplate for the figure
We will go through the notebook and collect the suggestions here