cwyoo wrote:Class projects correspondences in Probabilistic Graphical Models, Fall 2026. Please post your weekly proposed outcome here every Friday.
Week 1 Proposed Outcome: Initial Class Project IdeaDataset Link:
https://www.icpsr.umich.edu/web/NAHDAP/studies/37786For my class project, I am considering using longitudinal data from the Population Assessment of Tobacco and Health (PATH) Study. Since my research interests are in tobacco, nicotine, and cancer prevention, I would like to explore a topic that aligns with my research area and could potentially be developed further beyond the course.
My current idea is to examine whether patterns and transitions of tobacco and other substance use differ across birth cohorts, particularly Millennials and Generation Z. Rather than focusing only on cigarette and e-cigarette dual use, I am interested in broader patterns involving cigarette use, e-cigarette use, other tobacco products, polytobacco use, alcohol use, and cannabis use. Since PATH contains multiple waves of longitudinal data, I am interested in exploring whether these repeated observations could help us understand how these behaviors are interconnected and how they change over time.
Since I am still learning probabilistic graphical modeling, I am not yet sure which method would be most appropriate for addressing this question. I am particularly interested in learning whether a Bayesian network or another graphical modeling approach covered in this course could be useful for examining these longitudinal relationships and whether the patterns differ across birth cohorts.
Some methodological questions I am currently considering are:
1.What type of probabilistic graphical model would be most appropriate for repeated longitudinal observations from PATH?
2.How could information from multiple PATH waves be incorporated into such a model?
3.How could birth-cohort differences be examined while also considering differences in age and historical period?
4.Since PATH has a complex survey design with longitudinal and replicate weights, how should these weights be considered in a probabilistic graphical modeling analysis?
5.How should I determine a reasonable number of tobacco, alcohol, cannabis, psychosocial, and other variables to include without making the analysis overly complex?
For the next step, I plan to review the PATH documentation and identify relevant variables that are measured consistently across waves. I would appreciate your guidance on whether this would be an appropriate direction for the class project and which statistical method covered in the course may be most suitable for exploring this research question.