I know about the zip function (which will zip according to the shortest list) and zip_longest (which will zip according to the longest list), but how would I zip according to the first list, regardless of whether it's the longest or not?
For example:
Input: ['a', 'b', 'c'], [1, 2]
Output: [('a', 1), ('b', 2), ('c', None)]
But also:
Input: ['a', 'b'], [1, 2, 3]
Output: [('a', 1), ('b', 2)]
Do both of these functionalities exist in one function?
You can repurpose the "roughly equivalent" python code shown in the docs for itertools.zip_longest to make a generalized version that zips according to the length of the first argument:
from itertools import repeat
def zip_by_first(*args, fillvalue=None):
# zip_by_first('ABCD', 'xy', fillvalue='-') --> Ax By C- D-
# zip_by_first('ABC', 'xyzw', fillvalue='-') --> Ax By Cz
if not args:
return
iterators = [iter(it) for it in args]
while True:
values = []
for i, it in enumerate(iterators):
try:
value = next(it)
except StopIteration:
if i == 0:
return
iterators[i] = repeat(fillvalue)
value = fillvalue
values.append(value)
yield tuple(values)
You might be able to make some small improvements like caching repeat(fillvalue) or so. The issue with this implementation is that it's written in Python, while most of itertools uses a much faster C implementation. You can see the effects of this by comparing against Kelly Bundy's answer.
Here's another take, if the goal is readable, easy to understand code:
def zip_first(first, *rest, fillvalue=None):
rest = [iter(r) for r in rest]
for x in first:
yield x, *(next(r, fillvalue) for r in rest)
This uses the two-argument form of next() to return the fill value for all iterables that are exhausted.
For exactly two iterables, this can be simplified to
def zip_first(first, second, fillvalue=None):
second = iter(second)
for x in first:
yield x, next(second, fillvalue)