| name | tidyverse-patterns |
| description | Modern tidyverse patterns for R including pipes, joins, grouping, purrr, and stringr. Use when mentions "native pipe", "pipe nativo", "pipe operator", "operador pipe", "|>", "new pipe", "novo pipe", "is this the new pipe", "é o novo pipe", "join_by", "join_by()", "join syntax", "sintaxe de join", "relacionamento de join", ".by grouping", ".by", "per-operation grouping", "agrupamento por operação", "agrupamento por operacao", "across()", "pick()", "reframe()", "list_rbind", "list_rbind()", "list_cbind", "list_cbind()", "modern tidyverse", "tidyverse moderno", "latest tidyverse", "tidyverse atual", "dplyr 1.1", "dplyr 1.1 features", "recursos dplyr 1.1", "latest dplyr", "modern patterns", "padrões modernos", "padroes modernos", "usar pipe nativo", "use native pipe", "usar |>", "use |>", "agrupar por operação", "agrupar por operacao", "per-operation group", writing modern tidyverse R code with latest patterns, or implementing current dplyr 1.1+ features. ONLY FOR R - do NOT activate for pandas pipe, Python dplyr, or non-R languages. |
| version | 1.1.0 |
| user-invocable | false |
| allowed-tools | Read, Grep, Glob |
Modern Tidyverse Patterns
Best practices for modern tidyverse development with dplyr 1.1+ and R 4.3+
Core Principles
- Use modern tidyverse patterns - Prioritize dplyr 1.1+ features, native pipe, and current APIs
- Profile before optimizing - Use profvis and bench to identify real bottlenecks
- Write readable code first - Optimize only when necessary and after profiling
- Follow tidyverse style guide - Consistent naming, spacing, and structure
Pipe Usage (|> not %>%)
- Always use native pipe
|> instead of magrittr %>%
- R 4.3+ provides all needed features
data |>
filter(year >= 2020) |>
summarise(mean_value = mean(value))
data %>%
filter(year >= 2020) %>%
summarise(mean_value = mean(value))
Join Syntax (dplyr 1.1+)
- Use
join_by() instead of character vectors for joins
- Support for inequality, rolling, and overlap joins
transactions |>
inner_join(companies, by = join_by(company == id))
transactions |>
inner_join(companies, join_by(company == id, year >= since))
transactions |>
inner_join(companies, join_by(company == id, closest(year >= since)))
transactions |>
inner_join(companies, by = c("company" = "id"))
Multiple Match Handling
- Use
multiple and unmatched arguments for quality control
inner_join(x, y, by = join_by(id), multiple = "error")
inner_join(x, y, by = join_by(id), multiple = "all")
inner_join(x, y, by = join_by(id), unmatched = "error")
Data Masking and Tidy Selection
- Understand the difference between data masking and tidy selection
- Use
{{}} (embrace) for function arguments
- Use
.data[[]] for character vectors
my_summary <- function(data, group_var, summary_var) {
data |>
group_by({{ group_var }}) |>
summarise(mean_val = mean({{ summary_var }}))
}
for (var in names(mtcars)) {
mtcars |> count(.data[[var]]) |> print()
}
data |>
summarise(across({{ summary_vars }}, ~ mean(.x, na.rm = TRUE)))
Modern Grouping and Column Operations
- Use
.by for per-operation grouping (dplyr 1.1+)
- Use
pick() for column selection inside data-masking functions
- Use
across() for applying functions to multiple columns
- Use
reframe() for multi-row summaries
data |>
summarise(mean_value = mean(value), .by = category)
data |>
summarise(total = sum(revenue), .by = c(company, year))
data |>
summarise(
n_x_cols = ncol(pick(starts_with("x"))),
n_y_cols = ncol(pick(starts_with("y")))
)
data |>
summarise(across(where(is.numeric), mean, .names = "mean_{.col}"), .by = group)
data |>
reframe(quantiles = quantile(x, c(0.25, 0.5, 0.75)), .by = group)
data |>
group_by(category) |>
summarise(mean_value = mean(value)) |>
ungroup()
Modern purrr Patterns
- Use
map() |> list_rbind() instead of superseded map_dfr()
- Use
walk() for side effects (file writing, plotting)
- Use
in_parallel() for scaling across cores
models <- data_splits |>
map(\(split) train_model(split)) |>
list_rbind()
summaries <- data_list |>
map(\(df) get_summary_stats(df)) |>
list_cbind()
plots <- walk2(data_list, plot_names, \(df, name) {
p <- ggplot(df, aes(x, y)) + geom_point()
ggsave(name, p)
})
library(mirai)
daemons(4)
results <- large_datasets |>
map(in_parallel(expensive_computation))
daemons(0)
String Manipulation with stringr
- Use stringr over base R string functions
- Consistent
str_ prefix and string-first argument order
- Pipe-friendly and vectorized by design
text |>
str_to_lower() |>
str_trim() |>
str_replace_all("pattern", "replacement") |>
str_extract("\\d+")
str_detect(text, "pattern")
str_extract(text, "pattern")
str_replace_all(text, "a", "b")
str_split(text, ",")
str_length(text)
str_sub(text, 1, 5)
str_c("a", "b", "c")
str_glue("Hello {name}!")
str_pad(text, 10, "left")
str_wrap(text, width = 80)
str_to_lower(text)
str_to_upper(text)
str_to_title(text)
str_detect(text, fixed("$"))
str_detect(text, regex("\\d+"))
str_detect(text, coll("e", locale = "fr"))
grepl("pattern", text)
regmatches(text, regexpr(...))
gsub("a", "b", text)
Vectorization and Performance
result <- x + y
map_dbl(data, mean)
map_chr(data, class)
sapply(data, mean)
result <- numeric(length(x))
for(i in seq_along(x)) {
result[i] <- x[i] + y[i]
}
Common Anti-Patterns to Avoid
Legacy Patterns
data %>% function()
inner_join(x, y, by = c("a" = "b"))
sapply()
mutate(data, !!paste0("new_", var) := value)
Performance Anti-Patterns
result <- c()
for(i in 1:n) {
result <- c(result, compute(i))
}
result <- vector("list", n)
for(i in 1:n) {
result[[i]] <- compute(i)
}
result <- map(1:n, compute)
Migration from Old Patterns
From Base R to Modern Tidyverse
subset(data, condition) -> filter(data, condition)
data[order(data$x), ] -> arrange(data, x)
aggregate(x ~ y, data, mean) -> summarise(data, mean(x), .by = y)
sapply(x, f) -> map(x, f)
lapply(x, f) -> map(x, f)
grepl("pattern", text) -> str_detect(text, "pattern")
gsub("old", "new", text) -> str_replace_all(text, "old", "new")
substr(text, 1, 5) -> str_sub(text, 1, 5)
nchar(text) -> str_length(text)
strsplit(text, ",") -> str_split(text, ",")
paste0(a, b) -> str_c(a, b)
tolower(text) -> str_to_lower(text)
From Old to New Tidyverse Patterns
data %>% function() -> data |> function()
group_by(data, x) |>
summarise(mean(y)) |>
ungroup() -> summarise(data, mean(y), .by = x)
across(starts_with("x")) -> pick(starts_with("x"))
by = c("a" = "b") -> by = join_by(a == b)
summarise(data, x, .groups = "drop") -> reframe(data, x)
gather()/spread() -> pivot_longer()/pivot_wider()
separate(col, into = c("a", "b")) -> separate_wider_delim(col, delim = "_", names = c("a", "b"))
extract(col, into = "x", regex) -> separate_wider_regex(col, patterns = c(x = regex))
Superseded purrr Functions (purrr 1.0+)
map_dfr(x, f) -> map(x, f) |> list_rbind()
map_dfc(x, f) -> map(x, f) |> list_cbind()
map2_dfr(x, y, f) -> map2(x, y, f) |> list_rbind()
pmap_dfr(list, f) -> pmap(list, f) |> list_rbind()
imap_dfr(x, f) -> imap(x, f) |> list_rbind()
walk(x, write_file)
walk2(data, paths, write_csv)