]> Sergey Matveev's repositories - public-inbox.git/log
public-inbox.git
3 years agonntp: delimit Newsgroup: header with commas
Eric Wong [Tue, 3 Nov 2020 22:55:59 +0000 (22:55 +0000)]
nntp: delimit Newsgroup: header with commas

...instead of spaces.  This is specified in RFC 5536 3.1.4.

Include references to RFC 1036, 5536 and 5537 in our docs while
we're at it.

Reported-by: Andrey Melnikov <temnota.am@gmail.com>
Link: https://public-inbox.org/meta/CA+PODjpUN5Q4gBFQhAzUNuMasVEdmp9f=8Uo0Ej0mFumdSwi4w@mail.gmail.com/
3 years agotls: epollbit: account for miscellaneous OpenSSL errors
Eric Wong [Fri, 30 Oct 2020 02:13:58 +0000 (02:13 +0000)]
tls: epollbit: account for miscellaneous OpenSSL errors

Apparently they happen (triggered by my -imapd instance), so
bail out by closing the underlying socket rather than stopping
the event loop and daemon process.

3 years agoactually remove xt/eml_check_roundtrip.t
Eric Wong [Sat, 17 Oct 2020 21:27:06 +0000 (21:27 +0000)]
actually remove xt/eml_check_roundtrip.t

Fixes: 6550226296e9db79 ("xt: remove eml_check_roundtrip")
3 years agoxt: remove eml_check_roundtrip
Eric Wong [Sat, 17 Oct 2020 08:17:01 +0000 (08:17 +0000)]
xt: remove eml_check_roundtrip

If there's no body ({bdy} field), ->each_part set the {bdy}
field to "\n" and the ->as_string result afterwards is one
extra "\n" byte longer than the original.

It's not worth extra cycles in common ->each_part calls to
ensure 100% round-trip matches of header-only messages (which
are likely spam), especially when the only difference is a
trailing "\n".

3 years agogit: introduce async_wait_all
Eric Wong [Sat, 17 Oct 2020 08:04:24 +0000 (08:04 +0000)]
git: introduce async_wait_all

->cat_async and ->check_async may trigger each other (in future
callers) while waiting, so we need a unified method to ensure
both complete.  This doesn't affect current code, but allows us
to slightly simplify existing callers.

3 years agotmpfile: modernize to 5.10.1+, note O_APPEND workaround
Eric Wong [Fri, 16 Oct 2020 07:05:10 +0000 (07:05 +0000)]
tmpfile: modernize to 5.10.1+, note O_APPEND workaround

Once again we'll need O_APPEND on a temporary file, so note we
support it, here; since Perl 5.32 is way too new to depend on
our users having.

3 years agogit: async: loop inflight checks for nested callbacks
Eric Wong [Fri, 16 Oct 2020 06:59:34 +0000 (06:59 +0000)]
git: async: loop inflight checks for nested callbacks

We need to loop the inflight check for nested callback
invocations to ensure we don't clog the pipe that feeds
`git cat-file'.

This bug was obscured by the fact that we're already
accounting for 64-char git OIDs with SHA-256 in the
pipe space calculation; perhaps we shouldn't do that.

3 years agogit: *_async: support nested callback invocations
Eric Wong [Fri, 16 Oct 2020 06:59:33 +0000 (06:59 +0000)]
git: *_async: support nested callback invocations

For external indices, we'll need to support nested cat_async
invocations to deduplicate cross-posted messages.

Thus we need to ensure we do not clobber the {inflight*} queues
while stepping through and ensure {cat_rbuf} is stored before
invoking callbacks.

This fixes the ->cat_async-only case, but does not yet
account for the mix of ->check_async interspersed with
->cat_async calls, yet.  More work will be needed on that
front at a later date.

3 years agogit: ensure ->destroy clobbers check_async read buffer
Eric Wong [Fri, 16 Oct 2020 06:59:32 +0000 (06:59 +0000)]
git: ensure ->destroy clobbers check_async read buffer

It's currently not a problem as ->destroy doesn't
happen for no reason, we'll need to ensure future uses of
->destroy correctly discard the check_async buffer.

3 years agoinbox: add uidvalidity method
Eric Wong [Fri, 16 Oct 2020 06:59:31 +0000 (06:59 +0000)]
inbox: add uidvalidity method

This will make it easier to deal with ExtSearchIdx, which
won't have msgmap.

3 years agoscripts/dupe-finder: restore $dbh variable
Kyle Meyer [Fri, 16 Oct 2020 02:08:52 +0000 (22:08 -0400)]
scripts/dupe-finder: restore $dbh variable

When dupe-finder was switched from ->search->{over_ro} to ->over, the
database handle was dropped.  Restore it because a spot downstream
uses it.

Fixes: 73e3a6ed6e95adc6 (use more idiomatic internal API for ->over access)
3 years agoadmin: preserve config ordering of `--all' switch
Eric Wong [Tue, 13 Oct 2020 01:34:30 +0000 (01:34 +0000)]
admin: preserve config ordering of `--all' switch

When `--all' is passed to -index and similar commands, process
them in the same order as what is given in the config file.
This ensures predictable behavior so admins can ensure certain
inboxes see updated indices before others.  For (upcoming)
external indices, this will ensure stable Xref: ordering for
predictable caching/memoization by NNTP clients.

3 years agomanifest: favor Cpanel::JSON::XS
Eric Wong [Sat, 3 Oct 2020 23:12:46 +0000 (23:12 +0000)]
manifest: favor Cpanel::JSON::XS

JSON::MaybeXS already favors Cpanel::JSON::XS (and has for many
years, now).  Allow users to skip installing JSON::MaybeXS if
they want an XS-based JSON implementation.

3 years agov2writable: use "HEAD" to match v1 indexing behavior
Eric Wong [Tue, 29 Sep 2020 19:43:50 +0000 (19:43 +0000)]
v2writable: use "HEAD" to match v1 indexing behavior

Users may want to change the default branch used for git epochs
in v2 (v1 SearchIdx always used whatever "HEAD" pointed to).

3 years agosearchidx: index lower-case List-Id value
Eric Wong [Mon, 28 Sep 2020 05:15:19 +0000 (05:15 +0000)]
searchidx: index lower-case List-Id value

We don't want a List-Id value being confused with a Xapian
term prefix, here.

Followup-to: 8b06cda3a3af3f0e ("mda: match List-Id insensitively")
3 years agogcf2: improve error handling and do not ->fail on wbuf
Eric Wong [Sun, 27 Sep 2020 22:12:48 +0000 (22:12 +0000)]
gcf2: improve error handling and do not ->fail on wbuf

For historical reasons, both Danga::Socket::write and
PublicInbox::DS::write will return 0 when data is buffered;
so Gcf2Client must not call ->fail when DS::write returns 0.

We'll also improve robustness by recreating the entire
Gcf2Client object if it does die for other reasons, instead of
risking mismatched fields due to deferred close.

We also need to ensure we only get one EPOLLERR wakeup and
issue EPOLL_CTL_DEL if ->event_step is triggered by a dying
Gcf2 process, so always register the FD with EPOLLONESHOT.

3 years agods: add missing label for systems w/o EPOLLEXCLUSIVE
Eric Wong [Sun, 27 Sep 2020 21:26:08 +0000 (21:26 +0000)]
ds: add missing label for systems w/o EPOLLEXCLUSIVE

Oops :x

3 years agoimap: avoid raising exception if client disconnects
Eric Wong [Sat, 26 Sep 2020 08:08:37 +0000 (08:08 +0000)]
imap: avoid raising exception if client disconnects

This ought to save a few cycles if a client disconnects while
in the middle of a (UID) FETCH.  This avoids:

  Can't call method "git" on an undefined value at .../PublicInbox/IMAP.pm

errors in stderr.

3 years agoxt: add eml ->as_string round trip checker
Eric Wong [Thu, 24 Sep 2020 20:51:45 +0000 (20:51 +0000)]
xt: add eml ->as_string round trip checker

Unlike Email::MIME, PublicInbox::Eml::as_string should be able
to round trip from the Perl object to a raw scalar and back
without changes.

3 years agosearchidx: fix (undocumented) --skip-docdata handling
Eric Wong [Thu, 24 Sep 2020 10:13:39 +0000 (10:13 +0000)]
searchidx: fix (undocumented) --skip-docdata handling

This switch is still undocumented, but we can reduce the scope
of our Xapian docdata dependency by moving its only caller to
SearchIdx.  This reduces the amount of code loaded by read-only
code paths.

3 years agov2writable: drop outdated {unindex_range} check
Eric Wong [Tue, 22 Sep 2020 18:49:50 +0000 (18:49 +0000)]
v2writable: drop outdated {unindex_range} check

{unindex_range} only exists in the $sync state, nowadays, not the
V2Writable ($self) object.  $sync->{unindex_range} won't be
populated if $regen_max is zero, either, unless somebody is
injecting importable commits into an epoch history, in which
this change will result in no-op indexing doing no work.

3 years agoidxstack: fix comment about file_char
Eric Wong [Tue, 22 Sep 2020 18:49:49 +0000 (18:49 +0000)]
idxstack: fix comment about file_char

It's `d' for deletes, not `a'.

3 years agomda: match List-Id insensitively
Eric Wong [Mon, 21 Sep 2020 20:58:09 +0000 (20:58 +0000)]
mda: match List-Id insensitively

This follows -watch commit b70473ab8296d31ebb600adb4fa8fe0ac5935ca8
to match List-Id headers case-insensitively.

Reported-by: Konstantin Ryabitsev <konstantin@linuxfoundation.org>
Link: https://public-inbox.org/meta/20200921180152.uyqluod7qxbwqubo@chatter.i7.local/
3 years agomid: drop repeated ';' in mid_escape() regular expression
Kyle Meyer [Sun, 20 Sep 2020 18:13:58 +0000 (14:13 -0400)]
mid: drop repeated ';' in mid_escape() regular expression

3 years agodoc: post-1.6 updates, start 1.7
Eric Wong [Sat, 19 Sep 2020 21:42:01 +0000 (21:42 +0000)]
doc: post-1.6 updates, start 1.7

I should've dropped "PENDING" notes before the 1.6 release;
they're dropped now, and a note is added to remind my future
self to drop them before 1.7.

3 years agoconfig: warn on multiple values for some fields
Eric Wong [Sun, 20 Sep 2020 01:43:15 +0000 (01:43 +0000)]
config: warn on multiple values for some fields

Our code doesn't support multi-values for these, and having
unexpected arrays leads to unexpected results (e.g. showing
stuff like "ARRAY(0xDEADBEEFADD12E55)" in user interfaces).  So
warn and only use the last value (matching git-config(1)
behavior without `--get-all').

3 years agogcf2: wire up read-only daemons and rm -gcf2 script
Eric Wong [Sat, 19 Sep 2020 09:37:14 +0000 (09:37 +0000)]
gcf2: wire up read-only daemons and rm -gcf2 script

It seems easiest to have a singleton Gcf2Client client object
per daemon worker for all inboxes to use.  This reduces overall
FD usage from pipes.

The `public-inbox-gcf2' command + manpage are gone and a `$^X'
one-liner is used, instead.  This saves inodes for internal
commands and hopefully makes it easier to avoid mismatched
PERL5LIB include paths (as noticed during development :x).

We'll also make the existing cat-file process management
infrastructure more resilient to BOFHs on process killing
sprees (or in case our libgit2-based code fails on us).

(Rare) PublicInbox::WWW PSGI users NOT using public-inbox-httpd
won't automatically benefit from this change, and extra
configuration will be required (to be documented later).

3 years agogcf2: require git dir with OID
Eric Wong [Sat, 19 Sep 2020 09:37:13 +0000 (09:37 +0000)]
gcf2: require git dir with OID

This amortizes the cost of recreating PublicInbox::Gcf2 objects
when alternates change in v2 all.git.

3 years agogcf2*: more descriptive package descriptions
Eric Wong [Sat, 19 Sep 2020 09:37:12 +0000 (09:37 +0000)]
gcf2*: more descriptive package descriptions

Hopefully this allows others to more quickly figure out what's
going on.

3 years agogcf2: transparently retry on missing OID
Eric Wong [Sat, 19 Sep 2020 09:37:11 +0000 (09:37 +0000)]
gcf2: transparently retry on missing OID

Since we only get OIDs from trusted local data sources
(over.sqlite3), we can safely retry within the -gcf2 process
without worry about clients spamming us with requests for
invalid OIDs and triggering reopens.

3 years agoadd gcf2 client and executable script
Eric Wong [Sat, 19 Sep 2020 09:37:10 +0000 (09:37 +0000)]
add gcf2 client and executable script

This should be able to replace multiple `git cat-file' for blob
retrieval, but adjustments may be needed.

3 years agot/gcf2: test changes to alternates
Eric Wong [Sat, 19 Sep 2020 09:37:09 +0000 (09:37 +0000)]
t/gcf2: test changes to alternates

Calling ->add_alternate won't pick up new additions to
$OBJDIR/info/alternates, unfornately.  Thus v2 inboxes will
need to do something to invalidate Gcf2 objects.

3 years agogcf2: libgit2-based git cat-file alternative
Eric Wong [Sat, 19 Sep 2020 09:37:08 +0000 (09:37 +0000)]
gcf2: libgit2-based git cat-file alternative

Having tens of thousands of inboxes and associated git processes
won't work well, so we'll use libgit2 to access the object DB
directly.  We only care about OID lookups and won't need to rely
on per-repo revision names or paths.

The Git::Raw XS package won't be used since its manpages don't
promise a stable API.  Since we already use Inline::C and have
experience with I::C when it comes to compatibility, this only
introduces libgit2 itself as a source of new incompatibilities.

This also provides an excuse for me to writev(2) to reduce
syscalls, but liburing is on the horizon for next year.

3 years agogit_async_cat: inline + drop redundant batch_prepare call
Eric Wong [Thu, 17 Sep 2020 08:57:51 +0000 (08:57 +0000)]
git_async_cat: inline + drop redundant batch_prepare call

$git->cat_async already calls $git->batch_prepare iff needed, so
we can reduce subroutine calls and inline a one-off subroutine
to save some memory, here.

3 years agodoc: txt2pre: more manpage URLs
Eric Wong [Thu, 17 Sep 2020 21:25:22 +0000 (21:25 +0000)]
doc: txt2pre: more manpage URLs

We host our own -imapd manpage, and we started using a few more
git commands (fast-import for ages).  We'll also need to link to
manpages.debian.org and live with long URLs for a few
non-standard manpages in software we reference.

3 years agodoc: flow: include -imapd
Eric Wong [Thu, 17 Sep 2020 21:25:21 +0000 (21:25 +0000)]
doc: flow: include -imapd

It's another read-only daemon, and it may see more usage than
-nntpd as more users have IMAP support than NNTP.

3 years agot/indexlevels-mirror: fix improperly skipped test
Eric Wong [Wed, 16 Sep 2020 22:17:08 +0000 (22:17 +0000)]
t/indexlevels-mirror: fix improperly skipped test

Oops :x

3 years agopublic-inbox 1.6.0 v1.6.0
Eric Wong [Wed, 16 Sep 2020 20:02:20 +0000 (20:02 +0000)]
public-inbox 1.6.0

3 years agogit_async_cat: fix outdated comment
Eric Wong [Wed, 16 Sep 2020 07:18:53 +0000 (07:18 +0000)]
git_async_cat: fix outdated comment

We replaced Danga::Socket with PublicInbox::DS roughly a year
before GitAsyncCat was introduced into our git history.

3 years agowwwtext: link to public-inbox.org/meta archives
Eric Wong [Tue, 15 Sep 2020 19:48:50 +0000 (19:48 +0000)]
wwwtext: link to public-inbox.org/meta archives

Since we're advertising our address at meta@public-inbox.org,
we should advertise the archives, too.

3 years agowwwstream: link to cgit URLs for coderepo
Eric Wong [Tue, 15 Sep 2020 20:15:24 +0000 (20:15 +0000)]
wwwstream: link to cgit URLs for coderepo

Hopefully this reduces the ambiguity between code for the
project(s) using public-inbox and the code for public-inbox
itself.

3 years agotreewide: relax allow >=40 chars for git OID
Eric Wong [Tue, 15 Sep 2020 19:51:37 +0000 (19:51 +0000)]
treewide: relax allow >=40 chars for git OID

This will help with eventual git SHA-256 transitions.

3 years agomid: rename MID_MAX to ID_MAX
Eric Wong [Tue, 15 Sep 2020 19:51:36 +0000 (19:51 +0000)]
mid: rename MID_MAX to ID_MAX

It's only used for HTML anchors which we will need indefinitely.

3 years agoimap: quiet uninitialized variable warning on FETCH
Eric Wong [Tue, 15 Sep 2020 06:11:18 +0000 (06:11 +0000)]
imap: quiet uninitialized variable warning on FETCH

This was triggered by blindly trying to FETCH an MSN (not
"UID FETCH") on an empty dummy inbox.  It's harmless, and
probably triggered by a wayward client or misbehaving bot.

3 years agoci/deps: add Plack::Test::ExternalServer for devtest
Eric Wong [Mon, 14 Sep 2020 22:02:23 +0000 (22:02 +0000)]
ci/deps: add Plack::Test::ExternalServer for devtest

More of our Plack tests exercise public-inbox-httpd, nowadays;
and ExternalServer lets us test it easily alongside generic PSGI
stuff.

3 years agot/imapd.t: skip dependent test on failure
Eric Wong [Mon, 14 Sep 2020 21:29:40 +0000 (21:29 +0000)]
t/imapd.t: skip dependent test on failure

We don't want to cascade failures/warnings when something else
breaks.  There's likely more of these to be fixed as we
encounter them.

3 years agodoc: TODO and release notes updates ahead of 1.6
Eric Wong [Mon, 14 Sep 2020 06:39:36 +0000 (06:39 +0000)]
doc: TODO and release notes updates ahead of 1.6

Some more things have happened...

And drop some items which are too expensive to support,
such as automatic mirroring.

3 years agotests: consistently check for xapian-compact
Eric Wong [Mon, 14 Sep 2020 06:29:26 +0000 (06:29 +0000)]
tests: consistently check for xapian-compact

We may need to test against development versions of Xapian,
which may rely on setting `XAPIAN_COMPACT=xapian-compact-1.5'.
Ensure it's possible to do that.

And add a missing check in t/xcpdb-reshard.t, too.

3 years agosigfd: fix typos and scoping on systems w/o epoll+kqueue
Eric Wong [Mon, 14 Sep 2020 03:42:30 +0000 (03:42 +0000)]
sigfd: fix typos and scoping on systems w/o epoll+kqueue

Unfortunately, I'm not sure how easy catching these at
compile-time, is.  Prototypes do not seem to check these
at compile time when crossing packages (not even with
exported subroutines).

3 years agodoc: Add piem to list of clients
Kyle Meyer [Sun, 13 Sep 2020 19:31:55 +0000 (15:31 -0400)]
doc: Add piem to list of clients

3 years agonntp: share more code between art_lookup callers
Eric Wong [Fri, 11 Sep 2020 07:32:33 +0000 (07:32 +0000)]
nntp: share more code between art_lookup callers

This prepares us for future changes to improve scalability to
many inboxes.

3 years agot/nntpd: add test for the XPATH command
Eric Wong [Fri, 11 Sep 2020 07:32:32 +0000 (07:32 +0000)]
t/nntpd: add test for the XPATH command

It's only in RFC 2980 (not 977 or 3977), but Net::NNTP has
supported it since 2001, at least.  We'll be making changes
to avoid pathological behavior, so test it, first.

3 years agotreewide: avoid `goto &NAME' for tail recursion
Eric Wong [Fri, 11 Sep 2020 07:32:31 +0000 (07:32 +0000)]
treewide: avoid `goto &NAME' for tail recursion

While Perl implements tail recursion via `goto' which allows
avoiding warnings on deep recursion.  It doesn't (as of 5.28)
optimize the speed of such dispatches, though it may reduce
ephemeral memory usage.

Make the code less alien to hackers coming from other languages
by using normal subroutine dispatch.  It's actually slightly
faster in micro benchmarks due to the complexity of `goto &NAME'.

3 years agowwwstream: show init + index instructions for -V1, too
Eric Wong [Thu, 10 Sep 2020 05:46:04 +0000 (05:46 +0000)]
wwwstream: show init + index instructions for -V1, too

This should've always been there.  I'm not sure how widely
spread 1.0 and earlier releases were, but we'll keep documenting
the version requirement.

3 years agosolver: async blob retrieval for diff extraction
Eric Wong [Thu, 10 Sep 2020 01:51:53 +0000 (01:51 +0000)]
solver: async blob retrieval for diff extraction

Like the rest of the WWW code, public-inbox-httpd now uses
git_async_cat to retrieve blobs without blocking the event loop.
This improves fairness when git blobs are on slow storage and
allows us to take better advantage of SMP systems.

3 years agosolver: break apart inbox blob retrieval
Eric Wong [Wed, 9 Sep 2020 06:26:18 +0000 (06:26 +0000)]
solver: break apart inbox blob retrieval

To avoid hogging the event loop in public-inbox-httpd when
many candidate messages match, we'll separate the steps to
ensure fairness on slow storage.

3 years agosolver: check one git coderepo and inbox at a time
Eric Wong [Wed, 9 Sep 2020 06:26:17 +0000 (06:26 +0000)]
solver: check one git coderepo and inbox at a time

With public-inbox-httpd, this mitigates the effect of slow git
blob storage with multiple coderepos configured for an inbox.
It's still synchronous for now (and may need to remain that way
for ->last_check_err), but no longer monopolizes the event loop
when checking multiple coderepos.

We don't yet support multi-inbox scanning, yet; but this also
prepares us for a future where we do.

We'll also support >=40 char blob OIDs in preparation for future
git SHA-256 support, too.

3 years agowwwlisting: avoid hogging event loop
Eric Wong [Wed, 9 Sep 2020 06:26:16 +0000 (06:26 +0000)]
wwwlisting: avoid hogging event loop

By using the just-introduced ConfigIter class.
And make ManifestJsGz a subclass of it to reduce duplication.

3 years agoextmsg: prevent cross-inbox matches from hogging event loop
Eric Wong [Wed, 9 Sep 2020 06:26:15 +0000 (06:26 +0000)]
extmsg: prevent cross-inbox matches from hogging event loop

With many inboxes, checking multiple SQLite repos will be slow
and time-consuming, so ensure we can schedule it fairly between
multiple inboxes.

3 years agot/cgi.t: show stderr on failures
Eric Wong [Wed, 9 Sep 2020 06:26:14 +0000 (06:26 +0000)]
t/cgi.t: show stderr on failures

This helped me diagnose an error I would've introduced
in the next commit.

3 years agoconfig: split out iterator into separate object
Eric Wong [Wed, 9 Sep 2020 06:26:13 +0000 (06:26 +0000)]
config: split out iterator into separate object

We will need to allow simultaneous iterators on the same
config object, since we'll need this for ExtMsg, NNTPD,
WwwListing, NewsWWW, and other places.

3 years agoconfig: flatten each_inbox and iterate_start args
Eric Wong [Wed, 9 Sep 2020 06:26:12 +0000 (06:26 +0000)]
config: flatten each_inbox and iterate_start args

In Perl, we can simplify callers by passing a single array
all the way down the stack instead of a single array ref which
needs to be expanded every call.

3 years agowww: manifest.js.gz generation no longer hogs event loop
Eric Wong [Wed, 9 Sep 2020 06:26:11 +0000 (06:26 +0000)]
www: manifest.js.gz generation no longer hogs event loop

It's still as slow as before with hundreds/thousands of inboxes,
but at least it's fair.  Future changes will allow it to be
cached and memoized with persistent HTTP servers.

3 years agouse "\&" where possible when referring to subroutines
Eric Wong [Wed, 9 Sep 2020 06:26:10 +0000 (06:26 +0000)]
use "\&" where possible when referring to subroutines

"*foo" is ambiguous in that it may refer to a bareword file handle;
so we'll use it where we can without triggering warnings.

PublicInbox::TestCommon::run_script_exit required dropping the
prototype, however.  We'll also future-proof by dropping "use
warnings" in Cgit.pm and use the less-ambiguous "//=" in Inbox.pm
while we're in the area.

3 years agosolver: drop warnings, modernize use v5.10.1, use SEEK_SET
Eric Wong [Wed, 9 Sep 2020 06:26:09 +0000 (06:26 +0000)]
solver: drop warnings, modernize use v5.10.1, use SEEK_SET

With Perl upstream preparing to deprecate things, we'll move
towards only enabling warnings during development via shebang
and stop enabling them via "use".

We'll also favor "use v5.10.1" over the Perl 5.6-compatible "use
5.010_001", since our code base never worked on 5.6.

Finally, were also importing SEEK_SET without using it, just use it
for readability since we can't avoid loading Fcntl in other
places and it'll get constant-folded, anyways.

3 years agoxt/solver: test with public-inbox-httpd, too
Eric Wong [Wed, 9 Sep 2020 06:26:08 +0000 (06:26 +0000)]
xt/solver: test with public-inbox-httpd, too

We'll be making changes to solver to make it even fairer
to slow clients on slow storage.  Ensure we test with
public-inbox-httpd-specific codepaths, since the generic
PSGI code paths are rare in production use.

3 years agowwwtext: config comment improvements
Eric Wong [Wed, 9 Sep 2020 21:23:26 +0000 (21:23 +0000)]
wwwtext: config comment improvements

Use the full URL of the inbox being mirrored to reduce ambiguity
(instead of just the inbox name).

Using asymmetric quotes (e.g `foo') improves readability for me
in that it's more obvious when a quote begins and ends.  It also
lights up fewer pixels and reduces visual noise compared to
double-quotes.

We'll also reflow the `mainrepo' vs `inboxdir' comment slightly
to emphasize the word `instead'.

3 years agowwwtext: don't blindly quote "git clone" destination
Eric Wong [Wed, 9 Sep 2020 21:23:25 +0000 (21:23 +0000)]
wwwtext: don't blindly quote "git clone" destination

Save screen space and light up fewer pixels to reduce visual noise.

3 years agowwwtext: describe the use of `coderepo' entries
Eric Wong [Wed, 9 Sep 2020 21:23:24 +0000 (21:23 +0000)]
wwwtext: describe the use of `coderepo' entries

The `solver' feature is not very obvious, give potential
users a hint about it.

3 years agonntp: fix cross-newsgroup Message-ID lookups
Eric Wong [Thu, 10 Sep 2020 09:38:39 +0000 (09:38 +0000)]
nntp: fix cross-newsgroup Message-ID lookups

We cannot blindly use the selected newsgroup for
HEAD/ARTICLE/BODY requests using Message-ID, since
those commands look across all newsgroups; not just
the selected one (if any).

So stuff a reference to the Inbox object into $smsg.
We can reduce args passed into set_nntp_headers() and
msg_hdr_write(), too.

Fixes: 0e6ceff37fc38f28 ("nntp: support slow blob retrievals")
3 years agowwwstream: fix "Atom feed" link
Eric Wong [Wed, 9 Sep 2020 21:38:17 +0000 (21:38 +0000)]
wwwstream: fix "Atom feed" link

Oops, I wanted to stop escaping double-quotes with `qq()' but
used `q()' instead :x

Fixes: 2f61828fcb727e51 ("www: make mirror instructions more prominent")
3 years agocontrib/css: limit <a> coloring to links, only
Eric Wong [Wed, 9 Sep 2020 20:24:55 +0000 (20:24 +0000)]
contrib/css: limit <a> coloring to links, only

We don't want <a> tags without href= attributes to be colored,
since the `<a id=mirror>' tag in the HTML footer is intended
as an anchor destination for `<a href=#mirror>' link at the
top.

3 years agowww: make mirror instructions more prominent
Eric Wong [Tue, 8 Sep 2020 08:29:14 +0000 (08:29 +0000)]
www: make mirror instructions more prominent

In order to fight the misconception that public-inboxes are
centralized, anchor "#mirror" to the clone instructions and
place an emphasis on "mirror", not just cloning.

While we're at it, better describe multi-epoch -V2 inboxes,
since some users do not seem to realize epochs consist of
different data.

3 years agov2writable: reuse read-only shard counting code
Eric Wong [Wed, 2 Sep 2020 11:04:21 +0000 (11:04 +0000)]
v2writable: reuse read-only shard counting code

We'll also fix the read-only code to ensure we notice missing
Xapian shards, since gaps would throw off our expectation that
Xapian document IDs and NNTP article numbers are interchangeable.

3 years agooveridx: document column uses
Eric Wong [Wed, 2 Sep 2020 11:04:20 +0000 (11:04 +0000)]
overidx: document column uses

This may be useful for keeping our heads on straight dealing
with IMAP, NNTP, JMAP, etc.

3 years agowwwaltid: drop unused sqlite3_missing function
Eric Wong [Wed, 2 Sep 2020 11:04:19 +0000 (11:04 +0000)]
wwwaltid: drop unused sqlite3_missing function

It's inlined into the main function, which we'll shorten
slightly with the defined-or (`//') operator.  Also noticed
and fixed a mismatched HTML tag.

3 years agoimap: drop old, pre-Parse::RecDescent search parser
Eric Wong [Wed, 2 Sep 2020 11:04:18 +0000 (11:04 +0000)]
imap: drop old, pre-Parse::RecDescent search parser

We switched to Parse::RecDescent during development and left
some dead code behind.

3 years agosearch: remove {over_ro} field
Eric Wong [Wed, 2 Sep 2020 11:04:17 +0000 (11:04 +0000)]
search: remove {over_ro} field

Only inbox accesses the read-only {over}, now, instead of going
through ->search.  This simplifies our object graph and avoids
potentially redundant FDs and DB handles pointing to the same
over.sqlite3 file.

3 years agosearch: replace ->query with ->mset
Eric Wong [Wed, 2 Sep 2020 11:04:16 +0000 (11:04 +0000)]
search: replace ->query with ->mset

Nearly all of the search uses in the production code rely on
a Xapian mset iterator being returned (instead of an array
of $smsg objects).  So default to returning the mset and move
the burden of smsg array conversion into the test cases.

3 years agotests: add "use strict" and declare v5.10.1 compatibility
Eric Wong [Wed, 2 Sep 2020 11:04:15 +0000 (11:04 +0000)]
tests: add "use strict" and declare v5.10.1 compatibility

strict.pm helped me find a typo in an upcoming recent change, so
ensure we use it since it does more good than harm.  We'll also
take the opportunity here to declare v5.10.1 compatibility level
to future-proof against Perl incompatibilities.

3 years agosearch: remove special case for blank query
Eric Wong [Wed, 2 Sep 2020 11:04:14 +0000 (11:04 +0000)]
search: remove special case for blank query

The special case (if any) belongs at a higher-level,
and this is another step towards removing {over_ro}-dependence
in our Search object.

3 years agouse more idiomatic internal API for ->over access
Eric Wong [Wed, 2 Sep 2020 11:04:13 +0000 (11:04 +0000)]
use more idiomatic internal API for ->over access

{over_ro} being a part of the Search object is a historical
oddity which will go away, soon.  Lets start removing its use in
tests and rarely-used helper scripts.

3 years agodisambiguate OverIdx and Over by field name
Eric Wong [Wed, 2 Sep 2020 11:04:12 +0000 (11:04 +0000)]
disambiguate OverIdx and Over by field name

We'll use {oidx} as the common field name for the read-write
OverIdx, here, to disambiguate it from the read-only {over}
field.  This hopefully makes it clearer which code paths are
read-only and which are read-write.

3 years agomsgmap: note how we use ->created_at
Eric Wong [Wed, 2 Sep 2020 11:04:11 +0000 (11:04 +0000)]
msgmap: note how we use ->created_at

It'll likely be used in the future for JMAP, detached indices,
and maybe other things.

3 years agot/run: Perl future proofing
Eric Wong [Mon, 31 Aug 2020 23:41:56 +0000 (23:41 +0000)]
t/run: Perl future proofing

Bareword file handles outside of STD(IN|OUT|ERR) seem to be on
the chopping block for Perl 8.  We'll also "use v5.10.1" to
guard against future incompatibilities.

3 years agoinit+convert: create non-existing directory hierarchies
Eric Wong [Tue, 1 Sep 2020 01:15:07 +0000 (01:15 +0000)]
init+convert: create non-existing directory hierarchies

Following "git init" as an example, we'll create every parent
path up to the one specified, instead of attempting to continue
on when Cwd::abs_path returns `undef'.

3 years agodoc: remove B<> (bold) markup from the remaining POD
Eric Wong [Tue, 1 Sep 2020 01:15:06 +0000 (01:15 +0000)]
doc: remove B<> (bold) markup from the remaining POD

B<> decreases readability of the POD source and is of dubious
usefulness in the man page.

3 years agowatch: add --help/-h support
Eric Wong [Tue, 1 Sep 2020 01:15:05 +0000 (01:15 +0000)]
watch: add --help/-h support

And avoid unnecessary POD markup in the man page.

3 years agoconfig: use defined-or (//) in a few places
Eric Wong [Tue, 1 Sep 2020 01:15:04 +0000 (01:15 +0000)]
config: use defined-or (//) in a few places

Just some golfing to reduce scrolling and hopefully readability.

3 years agomda+learn: add --help / -h support
Eric Wong [Tue, 1 Sep 2020 01:15:03 +0000 (01:15 +0000)]
mda+learn: add --help / -h support

"use Getopt::Long" doesn't seem too slow on a hot page cache,
and it's probably used frequently enough to be in cache.

We'll also start reducing the amount of markup in the .pod and
favoring verbatim text in documentation for readability in
source form, since the bold text seems excessive.

3 years agodaemon: support --help/-h in -httpd/imapd/nntpd
Eric Wong [Tue, 1 Sep 2020 01:15:02 +0000 (01:15 +0000)]
daemon: support --help/-h in -httpd/imapd/nntpd

For consistency with other commands, though the
protocol-specific options should refer users to
the manpage.

3 years agoscript/*: fold $usage into $help, support `-h' instead of -?
Eric Wong [Tue, 1 Sep 2020 01:15:01 +0000 (01:15 +0000)]
script/*: fold $usage into $help, support `-h' instead of -?

`-h' doesn't conflict with anything, and some users (including
git users) may be more accustomed to using it rather than the
rarely-seen-outside-of-Getopt::Long `-?' switch.

We can also rely on the GetOptions() function to emit a proper
error message instead of just "bad command-line args".

3 years agoedit+purge: support `--help' and `-h' like other commands
Eric Wong [Tue, 1 Sep 2020 01:15:00 +0000 (01:15 +0000)]
edit+purge: support `--help' and `-h' like other commands

And while we're at it, note edit is *destructive* to encourage
reading the fine manual.

3 years agoadmin: improve minimum version text
Eric Wong [Tue, 1 Sep 2020 01:14:59 +0000 (01:14 +0000)]
admin: improve minimum version text

"inboxes 1 inboxes not supported by ..." was non-sensical.
Now it'll show "-V1 inbox not supported by ...", instead.

3 years agoscript/*: set executable bit on -learn and -imapd
Eric Wong [Tue, 1 Sep 2020 01:14:58 +0000 (01:14 +0000)]
script/*: set executable bit on -learn and -imapd

It's useful to mark they're meant to be executable, even
if the shebang is useless.

3 years agot/v2dupindex: test indexing mirrors with duplicate messages
Eric Wong [Tue, 1 Sep 2020 05:55:45 +0000 (05:55 +0000)]
t/v2dupindex: test indexing mirrors with duplicate messages

While it's not a known problem, our deduplicating logic may
change in the future; or a BOFH could be manually injecting
duplicate messages directly into the git epoch repositories.

Ensure indexing in mirrors doesn't break when there's
duplicates.  This is in preparation for detached indices
for multi-inbox search.

3 years agoindex: check for xapian-compact when using --compact
Eric Wong [Tue, 1 Sep 2020 16:54:31 +0000 (16:54 +0000)]
index: check for xapian-compact when using --compact

Otherwise, users may be frustrated to discover it missing
a long indexing run.

3 years agoreplace ParentPipe with EOFpipe
Eric Wong [Mon, 31 Aug 2020 04:41:40 +0000 (04:41 +0000)]
replace ParentPipe with EOFpipe

ParentPipe was a subset of EOFpipe, except EOFpipe correctly
accounts for theoretical(*) spurious wakeups on the pipe.

(*) AFAIK, spurious wakeups are/were more likely on TCP sockets
    due to checksum failures, something that's not a problem on
    local pipes.  We're also not sharing pipes like we do with
    listen sockets on accept(2), so there's no chance of another
    process grabbing bytes (unless we have bugs in our code).

3 years agods: avoid unnecessary timer for waitpid
Eric Wong [Mon, 31 Aug 2020 04:41:39 +0000 (04:41 +0000)]
ds: avoid unnecessary timer for waitpid

It doesn't seem necessary, since we won't call dwaitpid()
until we see an EOF.

3 years agowatch: use EOFpipe to reduce dwaitpid wakeups
Eric Wong [Mon, 31 Aug 2020 04:41:38 +0000 (04:41 +0000)]
watch: use EOFpipe to reduce dwaitpid wakeups

It's a bit inefficient to use a pipe, here.  However, using
dwaitpid() on a process that's not expected to exit soon is
also inefficient as it causes excessive wakeups as most of
our inbox-writing code expects synchronous waitpid().

This only affects -watch instances configured for NNTP and IMAP
clients.